Method for training object recognition model using spatial information and computing device for performing same
By using the illuminance information and image data in the spatial information, the neural network model is trained, and the problem of low object recognition accuracy in dark or illuminance differences is solved, achieving higher recognition accuracy and stability.
Patent Information
- Application Number
- CN202380061053.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-10-05
- Publication Date
- 2025-05-06
AI Technical Summary
In environments where total darkness or significant illuminance differences are present, the neural network model has reduced object recognition accuracy.
By using spatial information, including illuminance information and image data, a neural network model for object recognition is trained. The method involves obtaining illuminance information for multiple points in the space, matching image pairs for different time periods to calculate illuminance differences, and using these data to generate training data to train neural network models.
The accuracy of object recognition is improved, especially in environments with large illumination changes, and the stability and reliability of neural network models when identifying objects is enhanced.
Smart Images

Figure CN119948537A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method of training a neural network model for object recognition by using spatial information about a specific space and a computing device for performing the method. Background Art
[0002] Nowadays, neural network models for object recognition are widely used even in common electronic devices around us. For example, an electronic device that performs an operation by moving within a specific space (e.g., a house), such as a robot vacuum cleaner, can recognize surrounding objects by using a neural network model and perform an operation based on the recognition result.
[0003] However, there is a problem in that the object recognition accuracy of the neural network model is reduced when the illumination difference is significant because the lights are turned off, when the space where the object is placed is completely dark, or because only a specific area of the space is very bright compared to other areas.
[0004] For example, with regard to a robot vacuum cleaner, when the robot vacuum cleaner generally generates a map by performing spatial modeling of the interior of a house, the lights are turned on and the house is completely bright. However, when the robot vacuum cleaner actually performs an operation (cleaning), because the lights in the house are turned off, the house is completely dark, and because light enters only specific areas (e.g., windows), the illumination difference between areas is large, and therefore, the possibility of misidentifying an object is high.
[0005] The above information is presented as background information only to assist with an understanding of the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with respect to the present disclosure. Summary of the invention
[0006] Solution to the problem Aspects of the present disclosure are to at least solve the above-mentioned problems and / or disadvantages and to provide at least the advantages described below. Therefore, one aspect of the present disclosure is to provide a method for training a neural network model for object recognition by using spatial information about a specific space and a computing device for executing the method.
[0007] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.
[0008] According to one aspect of the present disclosure, there is provided a method for training an object recognition model by using spatial information. The method includes: obtaining spatial information including illumination information corresponding to a plurality of points in space, obtaining illumination information corresponding to at least one point among the plurality of points from the spatial information, obtaining training data by using the obtained illumination information and an image obtained by capturing the at least one point, and training a neural network model for object recognition by using the training data.
[0009] According to one aspect of the present disclosure, there is provided a computing device for executing a method for training an object recognition model by using spatial information. The computing device includes a memory storing a program for training a neural network model and at least one processor, the at least one processor being configured to execute the program to obtain spatial information including illumination information corresponding to a plurality of points in space, obtain illumination information corresponding to at least one of the plurality of points from the spatial information, obtain training data by using the obtained illumination information and an image obtained by capturing the at least one point, and train the neural network model for object recognition by using the training data.
[0010] According to an embodiment of the present disclosure, a non-transitory computer-readable recording medium may have stored therein a program for executing at least one embodiment of the method on a computer.
[0011] According to an embodiment of the present disclosure, a computer program is stored in a computer-readable recording medium so as to execute at least one embodiment of the method on a computer.
[0012] Other aspects, advantages, and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description which, in conjunction with the accompanying drawings, discloses various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which: Figure 1 is a diagram showing an environment for training a neural network model for object recognition by using spatial information according to an embodiment of the present disclosure; Figure 2 is a diagram showing a detailed configuration of a mobile terminal according to an embodiment of the present disclosure; Figure 3 is a diagram showing a detailed configuration of a robot vacuum cleaner according to an embodiment of the present disclosure; Figure 4a and Figure 4b is a diagram illustrating a map generated by a robot vacuum cleaner according to various embodiments of the present disclosure; Figure 5 is a diagram illustrating a method of transmitting a content image and a style image captured at the same point to a mobile terminal, performed by a robot vacuum cleaner according to an embodiment of the present disclosure; Figure 6 is a diagram illustrating a method of obtaining an illumination difference between two images included in an image pair according to an embodiment of the present disclosure; Figure 7 is a diagram illustrating a method of giving priorities to a plurality of image pairs and selecting image pairs to be used for generating training data according to the priorities, performed by a mobile terminal according to an embodiment of the present disclosure; Figure 8 and Fig. 9 is a diagram illustrating a process of synthesizing an object with a content image and then performing style transfer and generating training data according to various embodiments of the present disclosure; Fig.10a and Fig.10b is a diagram illustrating a process of training a neural network model for object recognition by using generated training data according to various embodiments of the present disclosure; Fig.11 is a diagram illustrating a method of selecting an object to be synthesized with a content image according to an embodiment of the present disclosure; Fig.12a and Figure 12b is a diagram illustrating a method of selecting a region of a content image to be synthesized with an object according to various embodiments of the present disclosure; and Fig.13 , Fig.14 , Fig.15 , Fig.16 , Fig.17 and Fig.18 is a flowchart illustrating a method of training an object recognition model by using spatial information according to various embodiments of the present disclosure.
[0014] Throughout the drawings, it should be noted that like reference numbers are used to depict the same or similar elements, features, and structures. DETAILED DESCRIPTION
[0015] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details should be considered as merely exemplary. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications of the various embodiments described herein may be made without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0016] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.
[0017] It will be understood that singular forms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.
[0018] Throughout this disclosure, the expression "at least one of a, b, or c" refers to only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.
[0019] In the following description of the present disclosure, descriptions of technologies that are well known in the art and not directly related to the present disclosure are omitted. This is to clearly convey the gist of the present disclosure by omitting unnecessary descriptions. The terms used herein are those defined based on the functions in the present disclosure and may vary according to the intention of the user or operator, precedents, etc. Therefore, the terms used herein should be defined based on the meaning of the terms and the description throughout the specification.
[0020] For the same reason, some elements in the accompanying drawings are exaggerated, omitted or schematically shown. In addition, the size of each element may not reflect its actual size substantially. In each of the accompanying drawings, the same or corresponding elements are represented by the same reference numerals.
[0021] With reference to the embodiments of the present disclosure described below in conjunction with the accompanying drawings, the advantages and features of the present disclosure and the methods for achieving the same will become apparent. However, the present disclosure can be embodied in many different forms and should not be construed as being limited to the embodiments of the present disclosure set forth herein. These embodiments of the present disclosure are provided so that the present disclosure will be thorough and complete, and the scope of the present disclosure will be fully conveyed to those of ordinary skill in the art. The embodiments of the present disclosure may be defined according to the claims. In the specification, the same reference numerals represent the same elements. When describing the present disclosure, detailed descriptions of related well-known functions or configurations that may obscure the subject matter of the present disclosure are omitted. The terms used herein are those defined based on the functions in the present disclosure, and may vary according to the intentions of the user or operator, precedents, etc. Therefore, the terms used herein should be defined based on the meaning of the terms and the description throughout the specification.
[0022] It will be understood that each box shown in the flowchart and the combination of boxes shown in the flowchart can be implemented by computer program instructions. These computer program instructions can be loaded into a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, and the instructions executed by the processor of the computer or other programmable data processing device can generate a device for implementing the functions specified in (one or more) flowchart boxes. These computer program instructions can also be stored in a computer-usable or computer-readable memory, which can instruct the computer or other programmable data processing device to act in a specific way, and the instructions stored in the computer-usable or computer-readable memory can produce a product including an instruction device that implements the functions specified in (one or more) flowchart boxes. The computer program instructions can also be installed on a computer or other programmable data processing device.
[0023] In addition, each frame shown in the flow chart may represent a module, a fragment or a portion of a code including one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that in the embodiments of the present disclosure, the functions mentioned in the frame may not occur in the order shown. For example, depending on the functions involved, two frames shown in succession may actually be executed substantially at the same time, or the frames may sometimes be executed in the reverse order.
[0024] The term "... unit" used in the embodiments of the present disclosure refers to a software or hardware component that performs certain tasks, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). However, the "... unit" is not meant to be limited to software or hardware. The "... unit" may be configured in an addressable storage medium or may be configured to operate one or more processors. In the embodiments of the present disclosure, the "... unit" may include components (such as software components, object-oriented software components, class components, and task components), processes, functions, properties, procedures, subroutines, program code segments, drivers, firmware, microcodes, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided in the components and "... units" may be combined into fewer components and "... units" or further separated into additional components and "... units". In addition, the "... unit" in the embodiments of the present disclosure may include one or more processors.
[0025] The embodiment of the present disclosure relates to a method for training a neural network model capable of improving the accuracy of object recognition by using spatial information. Before describing the specific embodiments of the present disclosure, the meanings of terms frequently used in this specification are defined.
[0026] In the present disclosure, "spatial information" may include various types of information related to the characteristics of a space. According to an embodiment of the present disclosure, the spatial information may include an image obtained by capturing a plurality of points included in a specific space. In addition, the spatial information may include an image generated based on information obtained at a plurality of points included in a specific space. The information obtained at the plurality of points may be a captured image obtained by a camera, or may be sensor data obtained by another sensor (such as lidar, etc.). In addition, the information obtained from the plurality of points may be information obtained by combining a plurality of sensor data obtained by a plurality of sensors. In addition, according to an embodiment of the present disclosure, the spatial information may include a map of a specific space. The map may include information such as the structure of the space, the relationship between the areas included in the space, the position of objects located in the space, and the like.
[0027] In the present disclosure, "illuminance information" may include various types of information related to the brightness of multiple points included in a space. According to an embodiment of the present disclosure, an image obtained by capturing multiple points included in a specific space may include illuminance information. Each captured image may indicate the brightness level when the capture was performed (how bright the captured point was). In addition, each captured image may also include metadata corresponding to the illuminance information. According to an embodiment of the present disclosure, a map of a specific space may also include illuminance information. For example, information about each of a plurality of areas included in the map (such as average illuminance) may be recorded. A specific method of configuring a map to include illuminance information is described below.
[0028] A "robotic mobile device" may refer to any type of device that performs various operations by moving automatically or according to a user's command. The robotic mobile device described in the present disclosure may capture the surrounding environment to perform operations, identify objects included in the captured image, and perform operations based on the object recognition results. Therefore, a neural network model for object recognition may be installed on the robotic mobile device. A representative example of a robotic mobile device is a robot vacuum cleaner, and it is assumed in the present disclosure that the robotic mobile device is a robot vacuum cleaner, but the method of training a neural network model and the trained neural network model according to the embodiments described in the present disclosure are not limited thereto and may be used in various types of robotic mobile devices. In addition, the method of training a neural network model and the trained neural network model according to the embodiments described in the present disclosure may be used by any type of device that needs to perform object recognition instead of the robotic mobile device.
[0029] "Style transfer" means an operation of transferring the style of a specific image to another image. In this regard, the "style" may include the color, hue, and brightness of the image. An image including the transferred style is called a "style image", and the image to which the style is transferred is defined as a "content image". An output image can be generated by performing style transfer and transferring the style included in the style image to the content image. In other words, when style transfer is performed by using the content image and the style image as input, an image (output image) in which the style of the style image is transferred to the content image can be output.
[0030] Hereinafter, a method for training an object recognition model by using spatial information and a computing device for executing the method according to an embodiment of the present disclosure are described with reference to the accompanying drawings.
[0031] Figure 1 is a diagram showing an environment for training a neural network model for object recognition by using spatial information according to an embodiment of the present disclosure. Figure 1 , the robot vacuum cleaner 200 may obtain spatial information by moving in a specific space and transmit the spatial information to the mobile terminal 100, and the mobile terminal 100 may train a neural network model by using the received spatial information and then transmit the trained neural network model to the robot vacuum cleaner 200. At this time, the spatial information obtained by the robot vacuum cleaner 200 may include illumination information about a plurality of points in the space, and the mobile terminal 100 may use the spatial information and the illumination information included therein when training the neural network model. In addition, according to an embodiment of the present disclosure, the robot vacuum cleaner 200 may train the neural network model by itself without transmitting the spatial information to the mobile terminal 100.
[0032] According to an embodiment of the present disclosure, it is assumed that when the robot vacuum cleaner 200 transmits a captured image to the mobile terminal 100, the mobile terminal 100 generates training data by considering an illumination difference between received images, and performs training on a neural network model by using the generated training data.
[0033] exist Figure 1 The reason why the mobile terminal 100 performs training on the neural network model instead of the robot vacuum cleaner 200 in the embodiments of the present disclosure is that the mobile terminal 100 generally includes a processor with higher performance than the robot vacuum cleaner 200. According to an embodiment of the present disclosure, instead of the mobile terminal 100, the robot vacuum cleaner 200 may directly perform training on the neural network model, or a higher performance computing device (e.g., a laptop computer, a desktop computer, a server, etc.) may perform training on the neural network model. For example, the operations described as being performed by the mobile terminal 100 in the embodiments of the present disclosure may be performed by various computing devices capable of communication and operation.
[0034] According to an embodiment of the present disclosure, a home Internet of Things (IoT) server that controls an IoT device such as the robot vacuum cleaner 200 exists at home, and the home IoT server may instead perform the above-described operations performed by the mobile terminal 100 .
[0035] In addition, according to an embodiment of the present disclosure, operations described in the present disclosure as being performed by the robot vacuum cleaner 200 (such as obtaining spatial information by moving in a specific space, sending the obtained spatial information to the mobile terminal 100, and identifying objects in the space) may also be performed by another electronic device (for example, a housekeeper robot, a pet robot, etc.) instead of the robot vacuum cleaner 200.
[0036] According to an embodiment of the present disclosure, a neural network model that is first trained may have been installed on the robot vacuum cleaner 200, and the mobile terminal 100 may update the neural network model installed on the robot vacuum cleaner 200 by performing additional training on the neural network model using a received image and then transmitting the neural network model to the robot vacuum cleaner 200. The subject of training the neural network model may be flexibly determined according to the performance and available resources of each device included in the system.
[0037] In addition, according to an embodiment of the present disclosure, although the trained neural network model is not installed in the robot vacuum cleaner 200, a new neural network model may be trained according to the embodiments described in the present disclosure and installed on the robot vacuum cleaner 200. In this regard, the training of the new neural network model may also be performed by any one of the mobile terminal 100, the robot vacuum cleaner 200, or another computing device.
[0038] According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may obtain spatial information including illumination information corresponding to a plurality of points in a space, and transmit the spatial information to the mobile terminal 100. Specifically, the robot vacuum cleaner 200 may obtain spatial information by capturing a plurality of points in a space during spatial modeling and by capturing a plurality of points in a space during operation (cleaning). At this time, an image obtained by capturing a plurality of points corresponds to the spatial information.
[0039] In addition, according to an embodiment of the present disclosure, the robot vacuum cleaner 200 may generate a map of the space by using captured images during space modeling. At this time, the generated map also corresponds to the space information.
[0040] According to an embodiment of the present disclosure, an image obtained by capturing a plurality of points by the robot vacuum cleaner 200 may include illuminance information of the captured points. The brightness of the captured points is expressed in the image and may correspond to the illuminance information of the points. Therefore, the robot vacuum cleaner 200 or the mobile terminal 100 may obtain illuminance information corresponding to the plurality of points from the captured image.
[0041] In addition, as will be described below, according to an embodiment of the present disclosure, an image pair can be obtained by matching two images obtained by capturing the same point at different time periods (e.g., during space modeling and during operation), and an illumination difference can be obtained between the two images included in the image pair. In this regard, the illumination difference also corresponds to illumination information.
[0042] Reference Figure 1 , the robot vacuum cleaner 200 may send images A and B captured during space modeling and images A' and B' captured during operation (cleaning) to the mobile terminal 100. The robot vacuum cleaner 200 may first perform modeling of the space (e.g., a house) in which the robot vacuum cleaner 200 is used before performing an operation. For example, when a user purchases the robot vacuum cleaner 200 and drives the robot vacuum cleaner 200 in a house for the first time, the robot vacuum cleaner 200 may generate a map of the house (space) by capturing and analyzing (e.g., object recognition, semantic segmentation, etc.) the captured images by moving back and forth in the house. Refer to the following Figure 4a and Figure 4b Describes the generated map.
[0043] According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may transmit images A and B captured at regular intervals (e.g., time or distance) to the mobile terminal 100 during the process of performing modeling of the space. The robot vacuum cleaner 200 may record at least one of a location where capturing is performed or a captured point in a path map during the space modeling, and perform capturing based on the path map during subsequent operations, thereby obtaining captured images of the same point at two different time periods.
[0044] When the modeling of the space is completed, the robot vacuum cleaner 200 may perform a cleaning operation. According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may transmit captured images A' and B' to the mobile terminal 100 during operation. The robot vacuum cleaner 200 may perform capture at a location recorded on a path map generated during space modeling (a location where an image transmitted to the mobile terminal 100 during space modeling is captured), and transmit the captured images A' and B' to the mobile terminal 100.
[0045] As described above, the robot vacuum cleaner 200 can perform capture of the same point during space modeling and during operation by using the generated map (spatial information) and transmit the captured image to the mobile terminal 100. The reason for training the neural network model by using the captured image of the same point during space modeling and operation is as follows.
[0046] Typically, the situation in which the robot vacuum cleaner 200 performs space modeling is likely to be a state in which the user stays at home (i.e., a state in which the lights are turned on even during the day or at night). However, the situation in which the robot vacuum cleaner 200 performs a cleaning operation is likely to be a state in which the house is dark when the lights are turned off. Therefore, because the illumination difference between the images collected by the robot vacuum cleaner 200 during space modeling and the images captured during actual operation is large, there is a high possibility that the robot vacuum cleaner 200 may not accurately recognize objects located in front (e.g., clothes, wires, etc.) during operation. In an embodiment of the present disclosure, in order to solve this problem, training data is generated by considering the illumination difference between the images captured during space modeling and the images captured during operation, and the training data is used to train a neural network model for object recognition.
[0047] In an embodiment of the present disclosure, images A and B captured by the robot vacuum cleaner 200 during space modeling are referred to as content images, and images A' and B' captured by the robot vacuum cleaner 200 during operation are referred to as style images. According to an embodiment of the present disclosure, a content image and a style image obtained by capturing the same point may be matched as an image pair. According to an embodiment of the present disclosure, the mobile terminal 100 may generate training data by performing style migration using content images A and B and style images A' and B', which will be described again below.
[0048] Reference Figure 1 In the following figures, for ease of description, only two content images A and B and two style images A′ and B′ are shown, but it is obvious that a greater number of content images and style images may be used.
[0049] Figure 2 is a diagram illustrating a detailed configuration of a mobile terminal according to an embodiment of the present disclosure.
[0050] Reference Figure 2According to an embodiment of the present disclosure, the mobile terminal 100 may include a communication interface 110, an input / output interface 120, a memory 130, and a processor 140. However, the components of the mobile terminal 100 are not limited to the above examples, and the mobile terminal 100 may include more or fewer components than the above components. In an embodiment of the present disclosure, some or all of the communication interface 110, the input / output interface 120, the memory 130, and the processor 140 may be implemented in the form of a single chip, and the processor 140 may include one or more processors.
[0051] The communication interface 110 is a component that transmits and receives signals (control commands and data, etc.) with an external device by wire or wirelessly, and may be configured to include a communication chipset that supports various communication protocols. The communication interface 110 may receive a signal from the outside and output the signal to the processor 140, or transmit a signal output from the processor 140 to the outside.
[0052] The input / output interface 120 may include an input interface (e.g., a touch screen, hard buttons, a microphone, etc.) for receiving control commands or information from a user and an output interface (e.g., a display panel, a speaker, etc.) for displaying an execution result of an operation or a status of the mobile terminal 100 under the control of the user.
[0053] The memory 130 is a component that stores various programs or data, and may include a storage medium such as a read-only memory (ROM), a random access memory (RAM), a hard disk, a compact disk read-only memory (CD-ROM), and a digital video disk (DVD), or a combination of storage media. The memory 130 may not exist alone, and may be included in the processor 140. The memory 130 may include a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory. The memory 130 may store a program for performing operations according to an embodiment of the present disclosure described below. The memory 130 may provide the stored data to the processor 140 according to a request of the processor 140.
[0054] The processor 140 is a component that controls a series of processes so that the mobile terminal 100 operates according to the embodiments of the present disclosure described below, and may include one or more processors. In this regard, the one or more processors may be a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP)), a graphics-specific processor (such as a graphics processing unit (GPU) or a visual processing unit (VPU)), or an artificial intelligence-specific processor (such as a digital processing unit (NPU)). For example, when the one or more processors are artificial intelligence-specific processors, the artificial intelligence-specific processor may be designed with a hardware structure specifically for processing a specific artificial intelligence model.
[0055] The processor 140 may record data to the memory 130 or read data stored in the memory 130, and specifically, execute a program stored in the memory 130 to process data according to a predefined operation rule or an artificial intelligence model. Therefore, the processor 140 may perform operations described in the following embodiments of the present disclosure, and unless otherwise specified, operations described in the following embodiments of the present disclosure as to be performed by the mobile terminal 100 may be performed by the processor 140.
[0056] Figure 3 is a diagram illustrating a configuration of a robot vacuum cleaner according to an embodiment of the present disclosure.
[0057] Reference Figure 3 , the robot vacuum cleaner 200 according to an embodiment of the present disclosure may include a communication interface 210, an input / output interface 220, a memory 230, a processor 240, a camera 250, a motor 260, and a battery 270. However, the components of the robot vacuum cleaner 200 are not limited to the above examples, and the robot vacuum cleaner 200 may include more or less components than the above components. In an embodiment of the present disclosure, some or all of the communication interface 210, the input / output interface 220, the memory 230, and the processor 240 may be implemented in a single chip form, and the processor 240 may include one or more processors.
[0058] The communication interface 210 is a component that transmits and receives signals (control commands and data, etc.) with an external device by wire or wireless, and may be configured to include a communication chipset that supports various communication protocols. The communication interface 210 may receive a signal from the outside and output the signal to the processor 240, or transmit a signal output from the processor 240 to the outside. According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may transmit an image captured by a camera 250 to be described below to the mobile terminal 100 through the communication interface 210.
[0059] The input / output interface 220 may include an input interface (e.g., a touch screen, hard buttons, a microphone, etc.) for receiving control commands or information from a user and an output interface (e.g., a display panel, a speaker, etc.) for displaying an execution result of an operation or a status of the robot vacuum cleaner 200 under the control of the user.
[0060] The memory 230 is a component that stores various programs or data, and may include a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory 230 may not exist alone, and may be included in the processor 240. The memory 230 may include a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory. The memory 230 may store a program for performing operations according to the embodiments of the present disclosure described below. The memory 230 may provide the stored data to the processor 240 according to a request of the processor 240.
[0061] The processor 240 is a component that controls a series of processes so that the robot vacuum cleaner 200 operates according to the embodiments of the present disclosure described below, and may include one or more processors. In this regard, the one or more processors may be a general-purpose processor (such as a CPU, AP, or DSP), a graphics-specific processor (such as a GPU or VPU), or an artificial intelligence-specific processor (such as an NPU). For example, when the one or more processors are artificial intelligence-specific processors, the artificial intelligence-specific processor may be designed with a hardware structure specifically for processing a specific artificial intelligence model.
[0062] The processor 240 may record data to the memory 230 or read data stored in the memory 230, and specifically, execute a program stored in the memory 230 to process data according to a predefined operation rule or an artificial intelligence model. Therefore, the processor 240 may perform operations described in the following embodiments of the present disclosure, and unless otherwise specified, operations described as being performed by the robot vacuum cleaner 200 in the following embodiments of the present disclosure may be performed by the processor 240.
[0063] The camera 250 is a component that captures the surrounding environment of the robot vacuum cleaner 200. The robot vacuum cleaner 200 may obtain spatial information from an image captured by using the camera 250. According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may obtain images of multiple points in the space by capturing the front relative to the moving direction using the camera 250. The processor 240 may generate a map of the space by performing spatial modeling using the obtained images. In addition, the robot vacuum cleaner 200 may transmit the obtained images to the mobile terminal 100 during operation.
[0064] The motor 260 is a component that provides power required for the robot vacuum cleaner 200 to perform a cleaning operation. According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may move in a space and perform a cleaning operation (eg, suction) due to a driving force provided by the motor 260.
[0065] The battery 270 may supply power to a component group included in the robot vacuum cleaner 200 .
[0066] Hereinafter, an embodiment of the present disclosure in which the robot vacuum cleaner 200 and the mobile terminal 100 train a neural network model for object recognition by using spatial information is described below.
[0067] First, a method of generating a map of a space through space modeling performed by the robot vacuum cleaner 200 is described. Figure 4a and Figure 4b is a diagram illustrating a map generated by a robot vacuum cleaner according to various embodiments of the present disclosure.
[0068] Reference Figure 4a , a map generated according to an embodiment of the present disclosure may include multiple layers.
[0069] The base layer may include information about the overall structure of the space. The base layer may also include information about the path that the robot vacuum cleaner 200 travels during the space modeling, and thus, the base layer may correspond to the above-mentioned path map. The locations where capture is performed during the space modeling may be recorded in the path map, which will be referred to below. Figure 5 describe.
[0070] The semantic layer may include information about the regions into which space is conceptually divided. Figure 4a In the illustrated embodiment of the present disclosure, the space is divided into two rooms (Room 1 and Room 2), and the semantic layer includes information about the areas divided into different rooms.
[0071] The first to third layers correspond to real-time layers, and thus, information included in these layers may be changed according to environmental changes in the space, and in some cases, some layers may be deleted or new layers may be added.
[0072] Reference Figure 4a , the positions of objects (such as furniture, etc.) located in the space are displayed on all layers of the first layer to the third layer, the objects with the lowest movement possibility are displayed on the first layer, and the objects with the highest movement possibility are displayed on the third layer. According to an embodiment of the present disclosure, the objects displayed on the first layer to the third layer may all be displayed on one layer, and the configuration of the layers may be configured in various ways according to the purpose of the task to be performed through the map.
[0073] According to an embodiment of the present disclosure, the robot vacuum cleaner 200 can determine the position of an object located in a space through a map, and can set a moving path during operation. In addition, when an object is moved, added, or deleted in the space, the robot vacuum cleaner 200 can update a layer including the object.
[0074] As described above, the robot vacuum cleaner 200 may update a map according to a task to be performed, and according to an embodiment of the present disclosure, illumination information may be stored in the map for training a neural network model.
[0075] Reference Figure 4b , the robot vacuum cleaner 200 may store the illumination information in the fourth and fifth layers as real-time layers. Figure 4a The fourth and fifth layers may be added while the first to third layers are added, or at least one of the first to third layers may be deleted and the fourth and fifth layers may be added. In this way, the change, addition or deletion of the layers included in the map may be performed as needed.
[0076] Reference Figure 4b , the fourth and fifth layers show areas where the illumination difference between two different time periods (e.g., during space modeling and during operation) is included in a specific range. The fourth layer shows areas with an illumination difference of 20 or more, and the fifth layer shows areas with an illumination difference of 50 or more. Figure 6 Describes a method for obtaining the difference in illumination between two different time periods.
[0077] According to an embodiment of the present disclosure, when generating training data, illuminance information stored in the fourth and fifth layers may be used. According to an embodiment of the present disclosure, when the mobile terminal 100 generates training data, the robot vacuum cleaner 200 may transmit a map to the mobile terminal 100, or transmit only some layers included in the map or illuminance information included in the map to the mobile terminal 100, and then the mobile terminal 100 may generate training data by using the map, the layers included in the map, or the illuminance information included in the map.
[0078] For example, when the image (content image) used to generate training data is an image obtained by capturing a point in an area displayed on the fourth layer (an area with an illumination difference of 20 or more), the mobile terminal 100 can change the style of the content image by performing style migration on the content image according to the first setting value. In this case, the first setting value can be determined according to the illumination information corresponding to the area displayed on the fourth layer (the illumination difference is 20 or more). In short, the mobile terminal 100 can obtain illumination information corresponding to the area including the point where the content image is captured from the map (the fourth and fifth layers), and change the style of the content image by using the obtained illumination information, thereby generating training data.
[0079] The following describes a process in which the mobile terminal 100 generates training data and trains a neural network model by using the content images A and B and the style images A′ and B′ received from the robot vacuum cleaner 200 .
[0080] Figure 5is a diagram illustrating a method of transmitting a content image and a style image captured at the same point to a mobile terminal, performed by a robot vacuum cleaner according to an embodiment of the present disclosure.
[0081] Reference Figure 5 , the robot vacuum cleaner 200 may record the position 501 at which the content image A is captured on the path map, while transmitting the content image A captured during the space modeling to the mobile terminal 100. Then, the robot vacuum cleaner 200 may perform capturing at the position 501 recorded on the path map while performing the cleaning operation, and transmit the captured style image A' to the mobile terminal 100.
[0082] In this way, the robot vacuum cleaner 200 can record the locations where content images A and B are captured on the path map generated during space modeling, and capture style images A' and B' at the same locations as content images A and B with reference to the path map while performing a cleaning operation, and send the captured style images A' and B' to the mobile terminal 100.
[0083] When receiving content images A and B and style images A' and B' from the robot vacuum cleaner 200, the mobile terminal 100 may obtain a plurality of image pairs by matching images having corresponding positions with each other. For example, the mobile terminal 100 may match content image A and style image A' as one image pair, and match content image B and style image B' as another image pair.
[0084] Figure 6 is a diagram illustrating a method of obtaining an illumination difference between two images included in an image pair according to an embodiment of the present disclosure. Figure 6 The process of the method shown in FIG. 1 may be performed by the mobile terminal 100 .
[0085] As described above, according to the embodiments of the present disclosure, a content image and a style image captured at the same location may be matched and processed as one image pair. Figure 6 A process in which the mobile terminal 100 calculates the illumination difference between the content image A and the style image A′ included in the image pair A and A′ is shown.
[0086] Reference Figure 6 , the mobile terminal 100 may convert each of the content image A and the style image A′ into a grayscale image and then calculate an average value of pixels for each region within each image.
[0087] According to an embodiment of the present disclosure, the mobile terminal 100 may divide the content image A and the style image A' into a plurality of regions (eg, a first local region to an Nth local region), and calculate an illumination difference between the content image A and the style image A' in each of the divided plurality of regions.
[0088] The process 610 in which the mobile terminal 100 calculates the illumination difference between the first local areas 611 and 612 is described as follows. The mobile terminal 100 calculates an average 152 of the values of the pixels included in the first local area 611 of the content image A. Because the content image A is converted into a grayscale image, the pixels included in the content image A may have values from 0 to 255. This is also applied to the style image A'. Subsequently, the mobile terminal 100 calculates an average 125 of the values of the pixels included in the first local area 612 of the style image A'. The mobile terminal 100 calculates the difference 27 of the average 152 and 125 of the pixel values between the content image A and the style image A' of the first local areas 611 and 612.
[0089] The process 620 in which the mobile terminal 100 calculates the illumination difference between the second partial areas 621 and 622 is described as follows. The mobile terminal 100 calculates the average 144 of the pixel values included in the second partial area 621 of the content image A. Subsequently, the mobile terminal 100 calculates the average 122 of the pixel values included in the second partial area 622 of the style image A', and calculates the difference 22 between the two averages 144 and 122.
[0090] The relative illumination difference of each local area is referred to as a “local illumination difference.” The relative illumination difference between the first local areas 611 and 612 becomes a first local illumination difference, and the relative illumination difference between the second local areas 621 and 622 becomes a second local illumination difference.
[0091] The mobile terminal 100 may calculate the relative illumination difference of the Nth local area (Nth local illumination difference) in the same manner, and then calculate the average of the first local illumination difference to the Nth local illumination difference, thereby obtaining the “global illumination difference” of the image pair A and A′. In other words, the value obtained by averaging the first local illumination difference to the Nth local illumination difference of the specific image pair becomes the global illumination difference.
[0092] Multiple local areas can be set to not overlap each other, but Figure 6 As shown, some areas may be arranged to overlap each other.
[0093] Figure 7 is a diagram illustrating a method of giving priorities to a plurality of image pairs and selecting image pairs to be used for generating training data according to the priorities, performed by a mobile terminal according to an embodiment of the present disclosure.
[0094] Assume that the mobile terminal 100 is based on the above reference Figure 6 The described method computes the global illumination difference for each of the image pairs A and A' and B and B'.
[0095] Reference Figure 7, it can be seen that the global illumination difference of the image pair A and A′ calculated by the mobile terminal 100 is 24, and the global illumination difference of the image pair B and B′ is 55.
[0096] According to an embodiment of the present disclosure, the mobile terminal 100 may give priority to a plurality of image pairs based on the global illumination difference. Figure 7 In the illustrated embodiment of the present disclosure, because the global illumination difference 55 of the image pair B and B' is higher than the global illumination difference 24 of the image pair A and A', the mobile terminal 100 gives priority 1 to the image pair B and B' and gives priority 2 to the image pair A and A'.
[0097] According to an embodiment of the present disclosure, the mobile terminal 100 may generate training data by using a preset number of image pairs having a high priority among a plurality of image pairs. Figure 7 In the illustrated embodiment of the present disclosure, assuming that the mobile terminal 100 generates training data by using only one of two image pairs having a higher priority, the mobile terminal 100 may generate training data by using the image pair B and B′.
[0098] In the embodiments of the present disclosure, for ease of explanation, it is assumed that there are two image pairs in total, and the mobile terminal 100 generates training data by using one image pair with a higher priority among the two image pairs. However, differently from this, the mobile terminal 100 may generate training data by using some image pairs with a higher priority among a large number of image pairs or image pairs with a higher priority percentage.
[0099] The mobile terminal 100 may synthesize the object with the content image B, and perform style transfer (transfer the style of the style image) on the content image synthesized with the object, so as to generate training data by using the selected image pair B and B'. In addition, according to an embodiment of the present disclosure, when the object is already included in the content image to be used when generating the training data, the mobile terminal 100 may perform only style transfer without synthesizing the object with the content image.
[0100] In the following, reference is made to Figure 8 and Fig. 9 Describes the process of performing object synthesis and style transfer, and references Fig.10a and Fig.10b Describes the process of performing (additional) training on a neural network model by using the generated training data. Fig.11 A method of selecting an object to be composited with a content image is described, and reference is made to Fig.12a and Figure 12b A method of selecting a region of a content image to be composited with an object is described.
[0101] Figure 8 and Fig. 9 is a diagram illustrating a process of synthesizing an object with a content image and then performing style transfer and generating training data according to various embodiments of the present disclosure. As described above, when an object has been included in a content image, it can be obtained from a reference Figure 8 and Fig. 9 Only the operation of compositing the object with the content image is omitted in the described process.
[0102] Reference Figure 8 , describing a process of performing semantic segmentation on a content image B included in a selected image pair B and B'.
[0103] Reference Figure 8 , when receiving the input of the content image B, the semantic segmentation model 800 may perform semantic segmentation to output a segmented image Bs divided into a non-floor area 801 and a floor area 802. The obtained segmented image Bs may be used when performing style transfer later. Fig. 9 Describe this.
[0104] The semantic segmentation model 800 may be a software configuration implemented when the processor 140 of the mobile terminal 100 executes a program stored in the memory 130. For example, operations described as to be performed by the semantic segmentation model 800 may actually be performed by the processor 140 of the mobile terminal 100.
[0105] Reference Fig. 9 , the mobile terminal 100 may generate a synthesized image Bo1 by synthesizing the object o1 with the content image B. At this time, it is assumed that the annotation "socks" has been assigned to the object o1. The method of selecting the object o1 to be synthesized with the content image B may be implemented in various ways. According to an embodiment of the present disclosure, the mobile terminal 100 may extract an object image stored in its own database (DB) or a DB of another device (e.g., a cloud server, etc.), and synthesize the object image with the content image B. In addition, according to an embodiment of the present disclosure, the mobile terminal 100 may synthesize an object with a low recognition accuracy (reliability) among the objects recognized when the robot vacuum cleaner 200 performs a cleaning operation, with the content image B. Referring to FIG. Fig.11 Describe this.
[0106] The mobile terminal 100 may synthesize the object o1 in any region of the content image B, but according to an embodiment of the present disclosure, the mobile terminal 100 may select which region of the content image B to synthesize with the object o1 based on the illumination difference of each region between the content image B and the style image B'. Fig.12a and Figure 12b Describe this.
[0107] Optionally, according to an embodiment of the present disclosure, as described above, when the content image already includes an object, the operation of synthesizing the object with the content image may be omitted, and the subsequent process may be performed. For example, this is when the robot vacuum cleaner 200 obtains the content image (such as Fig. 9 In the case where the content image is included in the content image from the beginning, it is necessary to perform annotation on the object to generate training data. According to an embodiment of the present disclosure, the mobile terminal 100 equipped with a relatively high-performance neural network model for object recognition can perform annotation compared to the robot vacuum cleaner 200.
[0108] According to an embodiment of the present disclosure, the mobile terminal 100 may obtain a training image Bo1' by inputting the synthesized image Bo1 and the style image B' into the style transfer model 900. In this regard, the mobile terminal 100 may Figure 8 The segmented image Bs obtained by the described process is input into the style transfer model 900 for style transfer.
[0109] according to Fig. 9 The disclosed embodiment shown, from Figure 8 The segmented image Bs generated from the content image B can be input not only as a segmented image corresponding to the synthetic image Bo1, but also as a segmented image corresponding to the style image B'. Because it is difficult to obtain accurate results when performing semantic segmentation due to the low overall illumination of the style image B', the segmented image Bs obtained by performing semantic segmentation on the content image B is usually used.
[0110] According to an embodiment of the present disclosure, the style transfer model 900 may output a training image Bo1' by transferring the style of the style image B' to the synthesized image Bo1. Since the training image Bo1' includes the object o1 synthesized with the content image B, the same annotation ("socks") assigned to the object o1 may be assigned to the training image Bo1'. Alternatively, as described above, according to an embodiment of the present disclosure, when the object is included in the content image from the beginning, the mobile terminal 100 may recognize the object included in the content image and assign the recognition result as the annotation of the training image Bo1'.
[0111] The style transfer model 900 may be a software configuration implemented when the processor 140 of the mobile terminal 100 executes a program stored in the memory 130. For example, operations described as being performed by the style transfer model 900 may actually be performed by the processor 140 of the mobile terminal 100.
[0112] Fig.10a and Fig.10bis a diagram illustrating a process of training a neural network model for object recognition by using generated training data according to various embodiments of the present disclosure. Fig.10a , the mobile terminal 100 performs additional training on the neural network model, and refers to Fig.10b , the robot vacuum cleaner 200 performs additional training on the neural network model.
[0113] When the mobile terminal 100 generates the training data (training image) Bo1 ′ according to the above-described process, the mobile terminal 100 or the robot vacuum cleaner 200 may perform additional training on the neural network model by using the generated training data Bo1 ′.
[0114] Reference Fig.10a , the mobile terminal 100 may additionally train and update the neural network model for object recognition by using the generated training image Bo1′, and then transmit the updated neural network model to the robot vacuum cleaner 200 to request an update. The robot vacuum cleaner 200 may update a previously installed neural network model by using the neural network model received from the mobile terminal 100.
[0115] According to an embodiment of the present disclosure, when the reliability (accuracy) of the neural network model installed in the robot vacuum cleaner 200 is lower than a certain reference as a result of recognizing an object, the mobile terminal 100 may perform additional training on the neural network model. Fig.11 Describe this.
[0116] According to an embodiment of the present disclosure, the neural network model updated by additional training may correspond to a specific point in space. In this case, the point corresponding to the neural network model updated by additional training may be a point at which an image used when generating training data for additional training is captured. Details of this aspect are as follows.
[0117] A neural network model additionally trained by using training data generated by using an image of a first point captured among a plurality of points in a space is referred to as a first neural network model, and a neural network model additionally trained by using training data generated by using an image of a second point captured among the plurality of points is referred to as a second neural network model, the first neural network model and the second neural network model corresponding to the first point and the second point, respectively. Therefore, according to an embodiment of the present disclosure, during operation (cleaning), the robot vacuum cleaner 200 may use the first neural network model when identifying an object at the first point, and may also use the second neural network model when identifying an object at the second point.
[0118] In addition, according to an embodiment of the present disclosure, when additional training of a neural network model is performed multiple times, some parameters included in the neural network model may be fixed, and only the remaining parameters may be changed each time the additional training is performed, and thus, the robot vacuum cleaner 200 may use the neural network model by fixing some parameters of the neural network model and only changing the remaining parameters for each point including an object to be recognized.
[0119] according to Fig.10b In the embodiment of the present disclosure shown, the robot vacuum cleaner 200 performs additional training on the neural network model instead of the mobile terminal 100. Generally, because the performance of the processor 140 installed in the mobile terminal 100 is better than that of the processor 240 installed in the robot vacuum cleaner 200, the mobile terminal 100 performs additional training on the neural network model unless there are special circumstances, such as Fig.10a As in the illustrated embodiment of the present disclosure, the robot vacuum cleaner 200 may perform additional training according to circumstances.
[0120] Reference Fig.10b , the mobile terminal 100 may transmit the generated training image Bo1 ′ to the robot vacuum cleaner 200 and request additional training of the neural network model, and the robot vacuum cleaner 200 may perform additional training of the neural network model by using the received training image Bo1 ′.
[0121] Fig.11 is a diagram illustrating a method of selecting an object to be synthesized with a content image according to an embodiment of the present disclosure.
[0122] Reference Fig.11 , the robot vacuum cleaner 200 captures the surrounding environment by performing a cleaning operation, and recognizes an object included in the captured image. Fig.11 In the illustrated embodiment of the present disclosure, the neural network model installed on the robot vacuum cleaner 200 identifies an object included in an image captured at a specific time as a “plastic bag.” At this time, the reliability of the object recognition result (inference result) output by the neural network model of the robot vacuum cleaner 200 is 0.6.
[0123] According to an embodiment of the present disclosure, the robot vacuum cleaner 200 may recognize various objects by performing a cleaning operation, transmit a captured image including an object to the mobile terminal 100 when reliability of a recognition result is lower than a preset reference, and request object recognition and annotation assignment.
[0124] In addition, according to an embodiment of the present disclosure, the robot vacuum cleaner 200 can recognize various objects by performing cleaning operations, and can determine that additional training of the neural network model is required when the reliability of the recognition result is lower than a preset reference, and perform a process for additional training (for example, sending a captured image for generating training data to the mobile terminal 100).
[0125] exist Fig.11 In the illustrated embodiment of the present disclosure, it is assumed that when the reliability of the object recognition result is 0.7 or less, the robot vacuum cleaner 200 determines that additional training is required or uses an object when generating training data. Because the reliability 0.6 of the result of recognizing the object o1 as a "plastic bag" is lower than the reference value 0.7, the robot vacuum cleaner 200 may determine that additional training is required and transmit a captured image including the object o1 to the mobile terminal 100 to request object recognition.
[0126] According to an embodiment of the present disclosure, the neural network model installed on the mobile terminal 100 recognizes the object o1 included in the captured image received from the robot vacuum cleaner 200 as “socks”, and outputs a result in which the reliability of the recognition result is 0.9.
[0127] Because the result of the object recognition ("socks") in the mobile terminal 100 shows high reliability, the mobile terminal 100 may segment and extract the object o1 from the captured image, and assign the annotation "socks" to the extracted object o1. The image of the object o1 with the assigned annotation may be used for synthesis with the content image B, as described above with reference to Fig. 9 described.
[0128] Fig.12a and Figure 12b is a diagram illustrating a method of selecting a region of a content image to be synthesized with an object according to various embodiments of the present disclosure.
[0129] According to an embodiment of the present disclosure, the mobile terminal 100 may calculate the illumination difference of each region for the content image B and the style image B' included in the same image pair. According to an embodiment of the present disclosure, the mobile terminal 100 may divide the content image B and the style image B' into a plurality of regions (e.g., a first local region to an Nth local region), and calculate the illumination difference between the content image B and the style image B' in each of the divided plurality of regions.
[0130] Fig.12a An example is shown in which the mobile terminal 100 calculates the illumination difference between the xth local areas 1201a and 1201b and the yth local areas 1202a and 1202b. According to an embodiment of the present disclosure, the mobile terminal 100 may convert each of the content image B and the style image B' into a grayscale image and then calculate the illumination difference of each area.
[0131] The process of calculating the illumination difference between the x-th local areas 1201a and 1201b by the mobile terminal 100 is described as follows. The mobile terminal 100 calculates the average 56 of the pixel values included in the x-th local area 1201a of the content image B. Subsequently, the mobile terminal 100 calculates the average 33 of the pixel values included in the x-th local area 1201b of the style image B'. The mobile terminal 100 calculates the difference 23 between the average 56 and 33 of the pixel values between the content image B and the style image B' of the x-th local areas 1201a and 1201b.
[0132] The process in which the mobile terminal 100 calculates the illumination difference between the y-th local areas 1202a and 1202b is described as follows. The mobile terminal 100 calculates the average 57 of the pixel values included in the y-th local area 1202a of the content image B. Subsequently, the mobile terminal 100 calculates the average 211 of the pixel values included in the y-th local area 1202b of the style image B'. The mobile terminal 100 calculates the difference 154 of the average 57 and 211 of the pixel values between the content image B and the style image B' of the y-th local areas 1202a and 1202b.
[0133] The mobile terminal 100 may calculate the illumination differences between the content image B and the style image B′ of the first to Nth local regions in the same manner, and select a region having the largest illumination difference as a region to be synthesized with the object.
[0134] As described above, the object is synthesized with the region having the largest illumination difference between the content image B and the style image B', thereby generating training data under conditions where object recognition is difficult, thereby improving the performance of the neural network model.
[0135] Figure 12b Shows that by combining the object with Fig.12a The process shown is an embodiment of the present disclosure of selecting regions to synthesize to generate training data.
[0136] Reference Figure 12b , the mobile terminal 100 may generate a synthesized image Bo2 by synthesizing the object o2 with the yth local area 1202a of the content image B. At this time, it is assumed that the annotation "pet dog excrement" has been assigned to the object o2. As described above, the object o2 to be synthesized with the content image B may be an object image stored in the DB of the mobile terminal 100 or the DB of another device (e.g., a cloud server, etc.), or may be an object with low recognition accuracy (reliability) among the objects recognized when the robot vacuum cleaner 200 performs a cleaning operation.
[0137] According to an embodiment of the present disclosure, the mobile terminal 100 may obtain the training image Bo2' by inputting the synthesized image Bo2 and the style image B' into the style transfer model 900. In this regard, the mobile terminal 100 may Figure 8 The segmented image Bs obtained by the described process is input into the style transfer model 900 for style transfer.
[0138] according to Figure 12b The embodiment of the present disclosure shown in FIG. Figure 8 The segmented image Bs generated from the content image B can be input not only as a segmented image corresponding to the synthetic image Bo2, but also as a segmented image corresponding to the style image B'. Because it is difficult to obtain accurate results when performing semantic segmentation due to the low overall illumination of the style image B', the segmented image Bs obtained by performing semantic segmentation on the content image B is usually used.
[0139] According to an embodiment of the present disclosure, the style transfer model 900 may output a training image Bo2' by transferring the style of the style image B' to the synthesized image Bo2. Because the training image Bo2' includes the object o2 synthesized with the content image B, the annotation ("pet dog excrement") assigned to the object o2 may be assigned to the training image Bo2'.
[0140] Fig.13 , Fig.14 , Fig.15 , Fig.16 , Fig.17 and Fig.18 is a flowchart illustrating a method for training an object recognition model by using spatial information according to an embodiment of the present disclosure. Figures 13 to 18 A method of training an object recognition model by using spatial information according to various embodiments of the present disclosure is described. Since the operations described below are performed by the mobile terminal 100 or the robot vacuum cleaner 200 described above, the description included in the embodiments of the present disclosure provided above may also be equally applicable even when omitted below.
[0141] Reference Fig.13 In operation 1301, the mobile terminal 100 may obtain space information including illumination information corresponding to a plurality of points in space. Detailed operations included in operation 1301 are described in detail in Fig.14 Shown in.
[0142] Reference Fig.14 , in operation 1401, the mobile terminal 100 may obtain images of a plurality of points captured during a first time range. According to an embodiment of the present disclosure, the mobile terminal 100 may receive images captured by the robot vacuum cleaner 200 during space modeling.
[0143] In operation 1402 , the mobile terminal 100 may generate a map of a space by using images captured during a first time range.
[0144] In operation 1403 , the mobile terminal 100 may obtain images of a plurality of points captured during the second time range. According to an embodiment of the present disclosure, the mobile terminal 100 may receive images captured by the robot vacuum cleaner 200 during a cleaning operation.
[0145] The point at which capture is performed during the first time range is recorded on the map and then referenced when capture is performed during the second time range, so that images captured during the first time range and images captured during the second time range can be matched with each other at the capture point.
[0146] Return again Fig.13 In operation 1302, the mobile terminal 100 may obtain illumination information corresponding to at least one of the plurality of points from the spatial information. Detailed operations included in operation 1302 are described in detail in Fig.15 is shown in .
[0147] Reference Fig.15 In operation 1501, the mobile terminal 100 may obtain a plurality of image pairs by matching two images having capture points corresponding to each other based on spatial information. In operation 1502, the mobile terminal 100 may obtain an illumination difference of at least one of the plurality of image pairs. In other words, the mobile terminal 100 may calculate an illumination difference between two images included in at least one of the plurality of image pairs.
[0148] Return again Fig.13 In operation 1303, the mobile terminal 100 may obtain training data by using the obtained illumination information and an image obtained by capturing at least one point. Detailed operations included in operation 1303 are described in detail in Fig.16 is shown in .
[0149] Reference Fig.16 In operation 1601, the mobile terminal 100 may select at least some of the image pairs based on the obtained illumination difference. According to an embodiment of the present disclosure, the mobile terminal 100 may assign priorities to the image pairs in ascending order of illumination difference and select a certain number of image pairs with high priorities.
[0150] In operation 1602, the mobile terminal 100 may generate training data by using the selected image pair. Detailed operations included in operation 1602 are described in detail in Fig.17 is shown in .
[0151] Reference Fig.17, in operation 1701, the mobile terminal 100 may select an object to be synthesized with at least one of the image pairs. According to an embodiment of the present disclosure, the mobile terminal 100 may select an object from its own DB or a DB of another device, or may select an object having a recognition accuracy (reliability) that does not satisfy a certain reference from among objects recognized when the robot vacuum cleaner 200 performs a cleaning operation.
[0152] In operation 1702, the mobile terminal 100 may synthesize the selected object with the content image included in the selected image pair. According to an embodiment of the present disclosure, the mobile terminal 100 may synthesize the object in an arbitrary region on the content image. Alternatively, according to an embodiment of the present disclosure, the mobile terminal 100 may select a region on the content image to be synthesized with the object based on the illumination difference of each region between the content image and the style image, and synthesize the object with the selected region. Fig.18 An embodiment of the present disclosure in this regard is shown in FIG.
[0153] Reference Fig.18 In operation 1801, the mobile terminal 100 may divide the content image and the style image included in the selected image pair into a plurality of regions. In operation 1802, the mobile terminal 100 obtains the illumination difference between the content image and the style image of each of the plurality of regions. In operation 1803, the mobile terminal 100 selects a region having the largest illumination difference from the plurality of regions included in the content image. In operation 1804, the mobile terminal 100 synthesizes the object with the selected region.
[0154] Return again Fig.17 , in operation 1703 , the mobile terminal 100 may generate training data by transferring the style of the style image included in the selected image pair to the content image synthesized with the object.
[0155] Return again Fig.13 , in operation 1304, the mobile terminal 100 may train a neural network model for object recognition by using the obtained training data.
[0156] According to the above-mentioned embodiments of the present disclosure, a neural network model for object recognition is trained by using spatial information, thereby improving object recognition accuracy by reflecting the special circumstances of each space (for example, situations where there are significant differences in illumination over time or situations where specific objects are often found).
[0157] According to an embodiment of the present disclosure, a method for training an object recognition model by using spatial information may include: obtaining spatial information including illumination information corresponding to multiple points in space, obtaining illumination information corresponding to at least one point among the multiple points from the spatial information, obtaining training data by using the obtained illumination information and an image obtained by capturing at least one point, and training a neural network model for object recognition by using the training data.
[0158] According to an embodiment of the present disclosure, obtaining spatial information may include obtaining images of multiple points captured during a first time range, generating a map of the space by using the images captured during the first time range, and obtaining images of multiple points captured during a second time range.
[0159] According to an embodiment of the present disclosure, a plurality of locations at which capture may be performed during a first time range are recorded on a map, and images captured during a second time range may be images captured at the plurality of locations recorded on the map during the second time range.
[0160] According to an embodiment of the present disclosure, obtaining illumination information may include obtaining a plurality of image pairs by matching two images having capture points corresponding to each other based on spatial information, and obtaining an illumination difference of at least one of the plurality of image pairs.
[0161] According to an embodiment of the present disclosure, obtaining the training data may include: selecting at least some of the plurality of image pairs based on the obtained illumination difference, and generating the training data by using the selected at least some of the image pairs.
[0162] According to an embodiment of the present disclosure, obtaining the training data may include: generating the training data by performing style transfer after synthesizing the object with at least one of the selected at least some image pairs.
[0163] According to an embodiment of the present disclosure, when an image captured during a first time range is referred to as a content image and an image captured during a second time range is referred to as a style image, generating training data by using at least some of the selected image pairs may include: selecting an object to be synthesized with at least one of the at least some of the selected image pairs, synthesizing the selected object with the content image included in the at least one image pair, and generating training data by transferring the style of the style image included in the at least one image pair to the content image synthesized with the object.
[0164] According to an embodiment of the present disclosure, training a neural network model may include: after first training the neural network model, additionally training the neural network model by using generated training data, and selecting an object to be synthesized with at least one image pair may include: as a result of recognizing multiple objects by using the first trained neural network model, selecting at least one object having a recognition accuracy lower than a preset reference.
[0165] According to an embodiment of the present disclosure, a result of an object synthesized with a content image recognized by using a neural network model having a higher recognition accuracy than a neural network model trained first may be annotated in training data.
[0166] According to an embodiment of the present disclosure, synthesizing a selected object with a content image may include: dividing each of a content image and a style image included in at least one image pair into multiple regions, obtaining an illumination difference between the content image and the style image for each of the multiple regions, selecting a region with a maximum illumination difference from the multiple regions included in the content image, and synthesizing the object with the selected region.
[0167] According to an embodiment of the present disclosure, selecting at least some of the plurality of image pairs may include assigning priorities to the plurality of image pairs in ascending order of the obtained illumination differences, and selecting a preset number of image pairs in ascending order of the priorities.
[0168] According to an embodiment of the present disclosure, obtaining the illumination difference may include: dividing the images included in the multiple image pairs into multiple regions, obtaining the pixel value difference between the two images in each region of the multiple regions as the local illumination difference, and obtaining the average of the local illumination differences with respect to each image as the global illumination difference, and the multiple regions may overlap with each other.
[0169] According to an embodiment of the present disclosure, the images captured during the first time range may be images obtained by capturing multiple points while the robot mobile device moves during modeling of the surrounding environment, and the images captured during the second time range may be images obtained by capturing multiple points while the robot mobile device moves during operation.
[0170] According to an embodiment of the present disclosure, the spatial information may be a map of the space, the illumination information corresponding to each of a plurality of areas included in the map may be stored in the map, and obtaining training data may include: determining an area to which at least one point in the map belongs, obtaining illumination information corresponding to the determined area from the map, and changing the style of an image obtained by capturing at least one point by using the obtained illumination information.
[0171] According to an embodiment of the present disclosure, for each of a plurality of points, at least some of the parameters included in the neural network model may be different from each other, and training of the neural network model may include additionally training the neural network model by using generated training data after first training the neural network model, and the additionally trained neural network model may correspond to a point at which an image was captured for generating the training data used during the additional training.
[0172] According to an embodiment of the present disclosure, a computing device for executing a method for training an object recognition model by using spatial information may include: a memory storing a program for training a neural network model, and at least one processor configured to execute the program to obtain spatial information including illumination information corresponding to multiple points in space, obtain illumination information corresponding to at least one point among the multiple points from the spatial information, obtain training data by using the obtained illumination information and an image obtained by capturing at least one point, and train the neural network model for object recognition by using the training data.
[0173] According to an embodiment of the present disclosure, when obtaining spatial information, at least one processor may obtain images of multiple points captured during a first time range, generate a map of the space by using the images captured during the first time range, and obtain images of multiple points captured during a second time range.
[0174] According to an embodiment of the present disclosure, when obtaining illumination information, at least one processor may obtain multiple image pairs by matching two images having corresponding capture points based on spatial information, and obtain an illumination difference of at least one of the multiple image pairs.
[0175] According to an embodiment of the present disclosure, when obtaining training data, at least one processor may select at least some image pairs from among a plurality of image pairs based on the obtained illumination differences, and generate training data by using the selected at least some image pairs.
[0176] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and is embodied in a computer-readable medium. The terms "application" and "program" used herein refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or a portion thereof, suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" may include any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard drive, a compact disk (CD), a digital video disk (DVD), or any other type of memory.
[0177] In addition, machine-readable storage media may be provided as non-transitory storage media. "Non-transitory storage media" are tangible devices and may exclude wired, wireless, optical or other communication links that send transient electrical signals or other signals. "Non-transitory storage media" do not distinguish between situations where data is semi-permanently stored in the storage medium and situations where data is temporarily stored. For example, "non-transitory storage media" may include a buffer that temporarily stores data. Computer-readable media can be any available media that can be accessed by a computer, and includes all volatile / non-volatile and removable / non-removable media. Computer-readable media include media that can permanently store data and media that can store and later rewrite data, such as rewritable optical disks or erasable memory devices.
[0178] According to an embodiment of the present disclosure, the method according to various disclosed embodiments may be provided by being included in a computer program product. The computer program product as a commodity may be traded between a seller and a buyer. The computer program product may be distributed in the form of a device-readable storage medium (e.g., a compact disk read-only memory (CD-ROM)), or may be distributed (e.g., downloaded or uploaded) directly and online through an application store or between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be at least temporarily stored in a device-readable storage medium, such as a memory of a manufacturer's server, an application store's server, or a relay server, or may be temporarily generated.
[0179] Although the present disclosure has been specifically shown and described with reference to the embodiments of the present disclosure, it will be understood that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure. For example, the described techniques may be performed in an order different from the described methods, and / or the described elements (such as systems, structures, devices, or circuits) may be combined or integrated in a form different from the described methods, or may be replaced or substituted by other elements or equivalents to achieve appropriate results. Therefore, the above embodiments should be considered only in a descriptive sense, and not for the purpose of limitation. For example, each component described as a single type may be implemented in a distributed manner, and similarly, the components described as distributed may be implemented in a combined form.
[0180] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A method for training an object recognition model by using spatial information, comprising: obtaining spatial information including illumination information corresponding to a plurality of points in the space; obtaining illumination information corresponding to at least one of the plurality of points from the spatial information; obtaining training data by using the obtained illumination information and an image obtained by capturing the at least one point; as well as A neural network model for object recognition is trained by using training data.
2. The method according to claim 1, wherein: Get spatial information, including: obtaining images of the plurality of points captured during a first time range; generating a map of the space by using images captured during the first time range; and Images of the plurality of points captured during a second time range are obtained.
3. The method according to any one of claims 1 and 2, in, A plurality of locations at which capture was performed during a first time range are recorded on the map, and Wherein, the images captured during the second time range are captured at the plurality of locations recorded on the map during the second time range.
4. The method according to any one of claims 1 to 3, wherein: Obtain illumination information, including: obtaining a plurality of image pairs by matching two images having capture points corresponding to each other based on spatial information; and An illumination difference of at least one image pair among the plurality of image pairs is obtained.
5. The method according to any one of claims 1 to 4, wherein: Obtain training data, including: selecting at least some of the plurality of image pairs based on the obtained illumination differences; and Training data is generated by using at least some of the selected image pairs.
6. The method according to any one of claims 1 to 5, wherein: Obtaining training data includes: generating the training data by performing style transfer after synthesizing the object with at least one of the selected at least some image pairs.
7. The method according to any one of claims 1 to 6, wherein: When images captured during a first time range are referred to as content images and images captured during a second time range are referred to as style images, generating training data by using at least some of the selected image pairs includes: selecting an object to be composited with at least one of the selected at least some of the image pairs; synthesizing the selected object with a content image included in the at least one image pair; and Training data is generated by transferring the style of the style image included in the at least one image pair to a content image synthesized with the object.
8. The method according to any one of claims 1 to 7, in, Training the neural network model includes: after first training the neural network model, additionally training the neural network model by using the generated training data, Wherein, selecting the object to be synthesized with the at least one image pair includes: as a result of recognizing a plurality of objects by using a first trained neural network model, selecting at least one object having a recognition accuracy lower than a preset reference.
9. The method according to any one of claims 1 to 8, wherein: Composite selected objects with content images, including: dividing each of the content image and the style image included in the at least one image pair into a plurality of regions; Obtaining an illumination difference between a content image and a style image in each of the plurality of regions; selecting a region having a maximum illumination difference from among the plurality of regions included in the content image; and Composites the object with the selected area.
10. The method according to any one of claims 1 to 9, wherein: Selecting at least some of the plurality of image pairs comprises: assigning priorities to the plurality of image pairs in ascending order of the obtained illumination differences; and A preset number of image pairs are selected in ascending order of priority. 11 . A non-transitory computer-readable recording medium having recorded thereon a program for executing the method according to claim 1 on a computer.
12. A computing device comprising: A memory storing a program for training a neural network model; as well as At least one processor is configured to execute the program to perform the following operations: obtaining spatial information including illumination information corresponding to a plurality of points in the space, obtaining illumination information corresponding to at least one of the plurality of points from the spatial information, obtaining training data by using the obtained illumination information and an image obtained by capturing the at least one point, and A neural network model for object recognition is trained by using training data.
13. The computing device of claim 12, wherein: The at least one processor is further configured to, when obtaining the spatial information: obtaining images of the plurality of points captured during a first time range, generating a map of the space by using images captured during the first time range, and Images of the plurality of points captured during a second time range are obtained.
14. The computing device according to any one of claims 12 and 13, wherein: The at least one processor is further configured to, when obtaining the illumination information: obtaining a plurality of image pairs by matching two images having capture points corresponding to each other based on spatial information, and An illumination difference of at least one image pair among the plurality of image pairs is obtained.
15. The computing device according to any one of claims 12 to 14, wherein: The at least one processor is further configured to, when obtaining the training data: selecting at least some of the plurality of image pairs based on the obtained illumination differences, and Training data is generated by using at least some of the selected image pairs.