Method and electronic device for training neural network model by enhancing images representing objects captured by multiple cameras

By using the conversion relationship between the first and second cameras in the IoT environment, the recognition results of the neural network model are converted, and the accuracy and stability of multi-view object image recognition are solved, and more efficient neural network model training and object recognition effects are achieved.

CN120112934APending Publication Date: 2025-06-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074957.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-06
Filing Date
2023-10-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the IoT environment, it is difficult for the prior art to effectively train neural network models to identify images of objects captured by multiple cameras, especially in image recognition at different perspectives, where recognition accuracy and stability problems exist.

Method used

By obtaining the object recognition results predicted by the first neural network model, and based on the conversion relationship with the first camera and the second camera, the conversion recognition results are converted to adapt to the second perspective, and training data is generated for training the second neural network model.

Benefits of technology

It improves the accuracy and stability of image recognition at different perspectives, enhances the representation ability of neural network models to multi-view object images, and improves object recognition tasks in IoT environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112934A_ABST
    Figure CN120112934A_ABST
Patent Text Reader

Abstract

There is provided a computer-implemented method of training a neural network model by enhancing images representing an object captured by a plurality of cameras, the method includes: obtaining a first object recognition result predicted by a first neural network model using, as an input, a first image captured by a first camera capturing a space including at least one object from a first perspective, converting the obtained first object recognition result based on a conversion relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to a second camera capturing an image of the space from a second angle of view, and converting the obtained first object recognition result based on the first object recognition result converted with respect to the second angle of view, training data is generated by performing marking on a second image corresponding to the first image, the second image being captured by a second camera, and a second neural network model is trained by using the generated training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and electronic device for training a neural network model by enhancing images representing objects captured by multiple cameras. Background Art

[0002] The Internet has evolved from a human-centric connected network through which humans generate and consume information to an Internet of Things (IoT) network that exchanges and processes information between distributed elements such as objects. Internet of Everything (IoE) technology has also emerged, in which big data processing technology, etc., is combined with IoT technology via connection to cloud servers. IoT technology can be applied to various fields such as smart home appliances, smart homes, smart buildings, or smart cities through the integration and combination of existing information technologies and various industries.

[0003] Electronic devices connected in an IoT environment can collect, generate, analyze or process data, share data with each other, and use the data for their tasks. Recently, with the rapid development of computer vision, various types of electronic devices that use neural network models to perform visual tasks have been developed. Therefore, interest in the connection between various types of electronic devices in an IoT environment is increasing. Summary of the invention

[0004] According to one aspect of the present disclosure, there is provided a computer-implemented method for training a neural network model by enhancing images representing objects captured by multiple cameras, the method comprising obtaining a first object recognition result predicted by a first neural network model using a first image captured by a first camera as an input, the first camera capturing a space including at least one object from a first perspective, the method comprising transforming the obtained first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to a second camera, the second camera capturing the space from a second perspective, the method comprising generating training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective, the second image being captured by the second camera, and the method comprising training a second neural network model by using the generated training data.

[0005] According to one aspect of the present disclosure, an electronic device includes a camera, a communication unit, a memory configured to store at least one instruction, and at least one processor operably connected to the camera, the communication unit and the memory, the at least one processor being configured to execute the at least one instruction to obtain, through the communication unit, a first object recognition result predicted by a first neural network model, the first neural network model using a first image captured from a first perspective by a camera of an external device that captures a space including at least one object as an input, the at least one processor being configured to execute the at least one instruction to transform the obtained first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the camera of the external device and a second camera coordinate system corresponding to the camera that captures the space from a second perspective, the at least one processor being configured to execute the at least one instruction to generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective, the second image being captured by the camera, and the at least one processor being configured to execute the at least one instruction to train the second neural network model by using the generated training data.

[0006] According to one aspect of the present disclosure, a cloud server includes a communication unit, a memory storing at least one instruction, and at least one processor operably connected to the communication unit and the memory, wherein the at least one processor is configured to execute at least one instruction to obtain a first object recognition result predicted by a first neural network model using a first image captured by a first camera as an input, the first camera capturing a space including at least one object from a first perspective, the at least one processor is configured to execute at least one instruction to obtain a second image captured by a second camera through the communication unit, the second camera capturing the space from a second perspective, the at least one processor is configured to execute at least one instruction to transform the obtained first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system. The at least one processor is configured to execute at least one instruction to generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective, and the at least one processor is configured to execute at least one instruction to train the second neural network model by using the generated training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which:

[0008] Figure 1 A home Internet of Things (IoT) environment in which an electronic device and an external device are connected to each other according to an embodiment of the present disclosure is shown;

[0009] Figure 2a and 2b shows a spatial diagram according to an embodiment of the present disclosure;

[0010] Figure 3a , Figure 3b , Figure 3c and Figure 3d A method of using layers constituting a spatial graph according to an embodiment of the present disclosure is shown;

[0011] Figure 4 A method for obtaining a spatial map according to an embodiment of the present disclosure is shown;

[0012] Figure 5 shows object recognition results of images captured from different viewing angles according to an embodiment of the present disclosure;

[0013] Figure 6 A process for generating training data for training a second neural network model by using a first object recognition result predicted by a first neural network model using a first image as input according to an embodiment of the present disclosure is shown;

[0014] Figure 7 An example of a first neural network model or a second neural network model according to an embodiment of the present disclosure is shown;

[0015] Figure 8 A process of converting a first object recognition result predicted from a first image to obtain a first object recognition result converted relative to a second viewing angle according to an embodiment of the present disclosure is shown;

[0016] Fig. 9 An example of generating training data for each of a plurality of second images captured from a second perspective according to an embodiment of the present disclosure is shown;

[0017] Fig.10 and Fig.11 An example of generating training data for each of a plurality of second images captured from a plurality of different second viewing angles according to an embodiment of the present disclosure is shown;

[0018] Fig.12 A process of training a second neural network model by using training data generated by using output of a first neural network model according to an embodiment of the present disclosure is shown;

[0019] Fig.13 A flowchart for describing a method for training a neural network model according to an embodiment of the present disclosure is shown;

[0020] Fig.14 An example of training a neural network model on an electronic device according to an embodiment of the present disclosure is shown;

[0021] Fig.15 An example of training a neural network model on a cloud server according to an embodiment of the present disclosure is shown;

[0022] Fig.16 An example of training a neural network model on a cloud server according to an embodiment of the present disclosure is shown;

[0023] Fig.17 An example is shown in which an electronic device according to an embodiment of the present disclosure performs a task by recognizing a user's gesture;

[0024] Fig.18 and Fig.19 shows a configuration of an electronic device according to an embodiment of the present disclosure; and

[0025] Fig. 20 The configuration of a server according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0026] The terms used herein will be briefly described, and then the disclosure will be described in detail. As used herein, the term "coupling" and its derivatives refer to any direct or indirect communication between two or more elements, whether or not these elements are in physical contact with each other. The terms "send", "receive" and "communication" and their derivatives cover both direct and indirect communication. The terms "include" and "comprise" and their derivatives refer to including but not limited to this. The term "or" is an inclusive term, meaning "and / or". The phrase "associated with ... " and its derivatives refer to including, being included in ..., interconnected with ..., including, being included in ..., connected to or with ... connected, coupled to or with ... coupled, can communicate with ..., collaborate with ..., interlace, juxtapose, be close to, be bound to or with ... bound, have, have ... property, have to or with ... relationship, etc. The term "controller" refers to any device, system or part thereof that controls at least one operation. Such a controller can be implemented with hardware or a combination of hardware and software and / or firmware. The function associated with any particular controller can be centralized or distributed, whether local or remote. The phrase "at least one of..." when used with a list of items means that different combinations of one or more of the listed items may be used, and only one item in the list may be needed. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C, and any variations thereof. The expression "at least one of a, b, or c" may indicate only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof. Similarly, the term "group" means one or more. Thus, the group of items may be a single item or a collection of two or more items.

[0027] Although the terms used herein are selected from commonly used terms that are currently widely used in consideration of their functions in the present disclosure, these terms may be different according to the intention of ordinary technicians in the field, precedents, or the emergence of new technologies. In addition, in some cases, there are also terms arbitrarily selected by the applicant, and in this case, their meanings will be defined in detail in the specification. Therefore, the terms used herein are not just designations of terms, but the terms are defined based on the meanings of the terms and contents throughout the present disclosure.

[0028] Singular expressions may also include plural meanings as long as they do not contradict the context. All terms used herein, including technical and scientific terms, may have the same meanings as those generally understood by those skilled in the art. In addition, although terms such as "first" or "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element.

[0029] Throughout the specification, as used herein, terms such as “…er(or)”, “…unit”, “…module” and the like mean a unit that performs at least one function or operation, which may be implemented as hardware or software or a combination thereof.

[0030] In addition, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or a portion thereof, suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard drive, a compact disk (CD), a digital video disk (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Non-transitory computer-readable media include media in which data can be permanently stored and media in which data can be stored and later rewritten, such as rewritable optical disks or erasable memory devices.

[0031] According to the present disclosure, functions related to artificial intelligence are performed by a processor and a memory. The processor may include one or more processors. In this case, the one or more processors may be a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP)), a dedicated graphics processor (such as a graphics processing unit (GPU) or a visual processing unit (VPU)), or a dedicated artificial intelligence processor (such as a neural processing unit (NPU)). One or more processors perform control to process input data according to predefined operating rules or an artificial intelligence model stored in a memory. Alternatively, in the case where one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processor may be designed with a hardware structure dedicated to processing a specific artificial intelligence model.

[0032] The predefined operating rules or artificial intelligence models may be generated via a training process. Here, generating via a training process may mean generating predefined operating rules or artificial intelligence models that are set to perform desired characteristics (or purposes) by training a basic artificial intelligence model using a learning algorithm that utilizes a large amount of training data. According to the present disclosure, the training process may be performed by the device itself on which the artificial intelligence is executed or by a separate server or system. Examples of learning algorithms may include supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning, but are not limited thereto.

[0033] The artificial intelligence model may include multiple neural network layers. Each neural network layer has multiple weight values, and the neural network arithmetic operation is performed by the arithmetic operation between the arithmetic operation result of the previous layer and the multiple weight values. The multiple weight values ​​in each neural network layer can be optimized by the result of training the artificial intelligence model. For example, the multiple weight values ​​can be refined to reduce or minimize the loss or cost value obtained by the artificial intelligence model during the training process. The artificial neural network may include, for example, a deep neural network (DNN), and may include, for example, a convolutional neural network (CNN), a deep neural network (DNN), a recursive neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recursive deep neural network (BRDNN), a deep Q network (DQN), etc., but is not limited thereto.

[0034] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to allow those skilled in the art to easily implement the embodiments of the present disclosure. However, the present disclosure may be implemented in many different forms and should not be construed as being limited to the embodiments of the present disclosure set forth herein.

[0035] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.

[0036] Figure 1is a diagram showing a home Internet of Things (IoT) environment in which an electronic device 100 and external devices are connected.

[0037] In an embodiment, the electronic device 100 is a robot vacuum cleaner. In some embodiments, the electronic device 100 may be various types of auxiliary robots for user convenience, mobile devices, augmented reality (AR) devices, virtual reality (VR) devices, or devices for detecting the surrounding environment and providing certain services in a specific location or space. The electronic device 100 may be equipped with various types of sensors for scanning a space (e.g., an area, a room, a room) and detecting objects in the space, as well as a neural network model. For example, the electronic device 100 may include at least one of an image sensor such as a camera, a light detection and ranging (LiDAR) sensor such as a laser distance sensor (LDS), or a time of flight (ToF) sensor. The electronic device 100 may be equipped with at least one model such as a DNN, a CNN, a RNN, or a BRDNN, and a combination thereof may be used.

[0038] The external devices connected to the electronic device 100 include a cloud server 200 and various types of IoT devices 300-1, 300-2, and 300-3. Figure 1 As shown, the IoT device may include a housekeeping robot 300-1, a pet robot 300-2, a smart home camera 300-3, etc., but is not limited thereto and may be of the same type as the electronic device 100. Each of the housekeeping robot 300-1, the pet robot 300-2, and the smart home camera 300-3 may scan a space and detect an object in the space by using various types of sensors.

[0039] According to an embodiment of the present disclosure, each of the electronic device 100, the housekeeper robot 300-1, the pet robot 300-2, and the smart home camera 300-3 can generate and store a space map as space information about a space including at least one object by using the collected space scanning information or object information. The electronic device 100, the housekeeper robot 300-1, the pet robot 300-2, and the smart home camera 300-3 can send or receive and store the space scanning information, the object information, or the space map to share the space scanning information, the object information, or the space map with each other.

[0040] Even when the devices are in the same space, they scan the space and detect objects from different angles and different viewing angles according to the position, performance or sensing range of each device. Whether each device is stationary or moving, the operating behavior of each device, etc. and thus the sensing information (such as images or audio) obtained by any one of the devices may be useful for training an artificial intelligence model equipped in another device.

[0041] According to an embodiment of the present disclosure, any one of the electronic device 100, the housekeeper robot 300-1, the pet robot 300-2, and the smart home camera 300-3 can be used as a master device or a server device, and the other devices can be used as slave devices or client devices. The device corresponding to the master device or the server device can receive, store, and manage spatial scanning information, object information, or spatial maps from other IoT devices. The device corresponding to the master device or the server device can classify, store, and manage the received information according to the location where the information is obtained. For example, the device corresponding to the master device or the server device can classify, collect, and manage such information according to whether the spatial scanning information, object information, or spatial map corresponds to the same space, zone, or area. The device corresponding to the master device or the server device can update the first information stored therein with the second information corresponding to the same location to maintain the latest and accuracy of the location-related information.

[0042] According to an embodiment of the present disclosure, the electronic device 100, the housekeeper robot 300-1, the pet robot 300-2, and the smart home camera 300-3 may send space scanning information, object information, or a space map to the cloud server 200 to store and manage the space scanning information, object information, or a space map through the cloud server 200. For example, when it is impossible to send the space scanning information, object information, or a space map to the electronic device 100 because the IoT device is powered off or is performing a specific function, the electronic device 100 may request and receive the space scanning information, object information, or a space map from the cloud server 200.

[0043] like Figure 1 As shown, the cloud server 200 can manage the spatial scanning information, object information or spatial map received from each of the electronic device 100, the housekeeper robot 300-1, the pet robot 300-2 and the smart home camera 300-3, and monitor the space in the home. The cloud server 200 can store and manage the spatial scanning information, object information or spatial map collected from multiple IoT devices for each registered user account or registered location. For example, the cloud server 200 can classify, collect and manage the spatial scanning information, object information or spatial map according to whether such information corresponds to the same space or zone. In response to a request from an electronic device 100, a housekeeper robot 300-1, a pet robot 300-2 or a smart home camera 300-3 located at home, the cloud server 200 can send information about the space in the home (such as a spatial map) to the corresponding device.

[0044] According to an embodiment of the present disclosure, an artificial intelligence (AI) hub (e.g., an AI speaker) located at home may receive and store and manage space scan information, object information, or space maps from IoT devices at home, instead of the cloud server 200. The AI ​​hub may store and manage space scan information, object information, or space maps collected from multiple IoT devices for each space or zone within the home.

[0045] According to an embodiment of the present disclosure, the AI ​​hub located at home can store and manage space scan information, object information, or space map together with the cloud server 200. For example, the AI ​​hub can process the space scan information or object information to generate or manage the space map, or convert the data to protect personal information and transmit the resultant data to the cloud server 200. The cloud server 200 can process the information received from the AI ​​hub to store and manage the space scan information, object information, or space map, and transmit them to the AI ​​hub.

[0046] According to an embodiment of the present disclosure, an electronic device 100 (e.g., a robot vacuum cleaner) can use a spatial map to perform tasks such as cleaning. To this end, the electronic device 100 can scan the space by using various types of sensors to update the spatial map with the latest spatial scanning information. The electronic device 100 can update the spatial map stored therein by using not only information sensed by itself but also part or all of the spatial map received from the cloud server 200, the housekeeper robot 300-1, the pet robot 300-2, and the smart home camera 300-3 connected in the home IoT environment.

[0047] For example, when a robot vacuum cleaner is fully charged at a charging station to clean a space in a home, the robot vacuum cleaner can perform cleaning by using a spatial map stored therein. The robot vacuum cleaner can use the most recently used spatial map to clean the same space. However, because the state of the space during the previous cleaning is different from the current state of the space, it is preferred to apply the latest information about the objects located in the space to the spatial map in order to perform effective cleaning. To this end, the robot vacuum cleaner may leave the charging station to travel the main route in advance and collect information about the objects in the space by itself. However, such pre-driving requires more time and battery consumption. In this case, the robot vacuum cleaner may receive the latest spatial map from another robot vacuum cleaner or at least one external device located in the same space to update the spatial map stored therein.

[0048] The robot vacuum cleaner may use part or all of a spatial map received from an external device. The robot vacuum cleaner may use a spatial map received from another robot vacuum cleaner of the same type, or may use information about an object whose position is expected to change frequently, to update a spatial map stored therein. Even when the robot vacuum cleaner receives a spatial map from a different type of device, the robot vacuum cleaner may use part or all of a spatial map for the same space to update a spatial map stored therein.

[0049] Figure 2a and 2b is a flowchart used to describe a spatial graph.

[0050] Figure 2a 1 shows a space map stored in the electronic device 100 as a robot vacuum cleaner and a hierarchical structure between a plurality of layers constituting the space map. Figure 2a As shown in , the spatial graph may include a base layer, a semantic layer, and a real-time layer, but is not limited thereto and may include more or fewer layers depending on the characteristics of the task to be performed.

[0051] The base layer provides information about the basic structure of the entire space, including walls, columns, and channels. By processing the three-dimensional point cloud data to match the coordinate systems with each other and store the positions, the base layer can provide three-dimensional information about the space, position information and trajectory information of objects, etc. The base layer is used as a base map and a geometric map.

[0052] The semantic layer provides semantic information on top of the base layer. The user of the electronic device 100 can assign semantic information such as "room 1", "room 2" or "restricted access area" to the basic structure of the entire space of the base layer, and use the semantic information to perform the task of the electronic device 100. For example, assuming that the electronic device 100 is a robot vacuum cleaner, the user can set semantic information in the semantic layer so that the robot vacuum cleaner only cleans 'room 2' or does not clean 'restricted access area'.

[0053] The real-time layer provides information about at least one object in the space. The at least one object may include a static object and a dynamic object. In the disclosure, the real-time layer may include a plurality of layers based on the object attribute information, and may have a hierarchical structure between the layers. Figure 2a As shown, the real-time layer may include a first layer, a second layer, and a third layer, but is not limited thereto, and may include more or less layers according to the standard of classification of the attribute information of the object. Figure 2a , the first layer includes system wardrobes and built-in cabinets, the second layer includes tables and sofas, and the third layer includes chairs.

[0054] Figure 2b Various examples of a real-time layer including multiple layers based on object attribute information are described.

[0055] The object attribute information can be used to classify the object according to objective criteria such as the object's type, shape, size, height, etc. or a combination of multiple criteria. In addition, because the object attribute information can vary depending on the user and the environment, the attribute information of each object can be marked and input in advance.

[0056] According to an embodiment of the present disclosure, when the object attribute information is the mobility level (ML) of the object, the first layer may include objects corresponding to ML 1, the second layer may include objects corresponding to ML 2 and ML 3, and the third layer may include objects corresponding to ML 4. The mobility level of the object may be determined by applying objective characteristics of the object to a predefined classification standard for evaluating mobility. For example, ML 1 corresponds to an immovable object, ML 2 corresponds to an object that is movable but mostly remains stationary, ML 3 corresponds to an object that is movable but moves occasionally, and ML 4 corresponds to an object that is movable and moves frequently.

[0057] According to an embodiment of the present disclosure, when the object attribute information is the position movement cycle of the object, the first layer may include objects that have not moved within a month, the second layer may include objects that have moved within a month, and the third layer may include objects that have moved within a week. Unlike the mobility level classified based on the object characteristics of the object, the position movement cycle of the same object may vary depending on the user of the object or the environment in which the object is located. For example, object 'A' may be frequently used by a first user, but rarely used by a second user. Object 'B' may be frequently used in a first location, but rarely used in a second location.

[0058] According to an embodiment of the present disclosure, when the object attribute information is the height at which the object is located, the first layer may include objects corresponding to less than 1m, the second layer may include objects corresponding to 1m to 2m, and the third layer may include objects corresponding to greater than 2m.

[0059] According to an embodiment of the present disclosure, the classification criteria of multiple layers included in the real-time layer can be defined by the user. For example, the user can generate a spatial diagram reflecting the characteristics of the task by combining and setting the attribute information of multiple types of objects for the classification criteria. For example, for a robot vacuum cleaner that usually moves below a height of 50 cm, there is no need to consider objects located above 1m, such as lamps or photo frames on the wall. Therefore, the user can manually set the criteria for classification of the layers so that the first layer includes objects corresponding to ML 1 and located at 1m or less, the second layer includes objects corresponding to ML 2 or ML 3 and located at 1m or less, and the third layer includes objects corresponding to ML 4 and located at 1m or less.

[0060] Figure 3a , Figure 3b , Figure 3c and Figure 3d It is a diagram for describing a method of using layers constituting a spatial graph.

[0061] Depending on the type of electronic device 100 and IoT device or the characteristics of the task, the spatial maps used by each device may be different from each other. The electronic device 100 may use the existing spatial map stored therein as is. However, when there is a change in the space where the task is to be performed, the electronic device 100 may update the spatial map to reflect the change. The electronic device 100 may update the existing spatial map by receiving a spatial map that already reflects the spatial change from at least one external device. The electronic device 100 may generate a new spatial map based on the existing spatial map.

[0062] refer to Figure 3a , the electronic device 100 may load an existing spatial map (hereinafter, "first spatial map") stored therein. The first spatial map includes a base layer, a first layer, a second layer, and a third layer. Hereinafter, the first layer to the third layer include Figure 2b In the case where the first spatial map has been generated only a few minutes ago or the space has not changed after the first spatial map was previously used, the electronic device 100 may use the first spatial map as is to obtain a new spatial map (hereinafter, referred to as a second spatial map) and use the second spatial map to perform a new task.

[0063] refer to Figure 3b , the electronic device 100 may load the first spatial map stored therein. In the case where the task to be performed by the electronic device 100 does not require information about frequently moving objects with ML 4 or only requires information about objects that have not moved for a week or more, the electronic device 100 may obtain the second spatial map by selecting a base layer, a first layer, and a second layer from the layers constituting the first spatial map or removing the third layer from the first spatial map.

[0064] refer to Figure 3c , the electronic device 100 may load the first spatial map stored therein. In the case where a new task to be performed by the electronic device 100 requires information about an object having ML 1 or only requires information about an object that has not moved for more than one month or more, the electronic device 100 may obtain a second spatial map by selecting a base layer and a first layer from the layers constituting the first spatial map or removing the second layer and the third layer from the first spatial map.

[0065] refer to Figure 3d, the electronic device 100 may load the first spatial map stored therein. In the case where a new task to be performed by the electronic device 100 needs to reflect the latest information about objects having movable ML 2, ML 3, and ML 4, the electronic device 100 may obtain a second spatial map by selecting a base layer and a first layer from the layers constituting the first spatial map or removing the second layer and the third layer from the first spatial map. Thereafter, the electronic device 100 may obtain a third spatial map by extracting the second layer and the third layer from the spatial map received from the external device and applying them to the second spatial map. Alternatively, the electronic device 100 may obtain a third spatial map by detecting objects corresponding to ML 2, ML 3, and ML 4 using at least one sensor provided in the electronic device 100 and reflecting them in the second spatial map.

[0066] Figure 4 is a flowchart of a method for obtaining a spatial map according to an embodiment of the present disclosure.

[0067] The electronic device 100 may obtain a first spatial map (S410). The first spatial map may include a plurality of layers based on object property information. The first spatial map may be generated by the electronic device 100, or may be received from a device outside the electronic device 100.

[0068] The electronic device 100 may determine whether it is necessary to update the first spatial map (S420). For example, the electronic device 100 may determine whether it is necessary to update the first spatial map according to the characteristics of the task. The "task" is set by the electronic device 100 to be performed by the sole purpose of the electronic device 100 or the function of the electronic device 100. The setting information related to the execution of the task may be manually input to the electronic device 100 by the user, or sent to the electronic device 100 through a terminal such as a mobile device or a dedicated remote controller. For example, in the case where the electronic device 100 is a robot vacuum cleaner, the task of the robot vacuum cleaner may be cleaning a room or an area set by the user, scheduled cleaning by a scheduling function, low noise mode cleaning, etc. When the information to be used for performing the task is insufficient, the electronic device 100 may determine that the first spatial map needs to be updated. When the latest information about the objects in the space where the task is to be performed is necessary, the electronic device 100 may determine that the first spatial map needs to be updated. In an embodiment, the electronic device 100 may determine whether it is necessary to update the first spatial map according to the time elapsed from the time when the first spatial map is obtained or the set update cycle. When there is no need to update the first spatial map, the electronic device 100 may use the first spatial map as the second spatial map for performing the task.

[0069] When the first spatial map needs to be updated, the electronic device 100 may obtain object information (S430). The electronic device 100 may collect spatial scanning information or object information by itself by using at least one sensor. The electronic device 100 may receive part or all of the spatial map or spatial scanning information or object information from an external device.

[0070] The electronic device 100 may update the first spatial map using the acquired spatial scanning information or object information by using the spatial scanning information or object information (S440). For example, for a frequently moving object, the electronic device 100 may newly acquire information about the object, or spatial scanning information about the location where the object is located, to update the first spatial map to reflect the latest location information.

[0071] The electronic device 100 may obtain the second spatial map (S450). The electronic device 100 may obtain the second spatial map by using the first spatial map as it is, by using the first spatial map in which some object information or some layers are modified, or by updating the first spatial map.

[0072] In an embodiment, the second spatial map to be used for performing the task may be converted into or generated in a map of an appropriate type according to the function of the electronic device 100 or the characteristics of the task, and then used. For example, in the case where the electronic device 100 is a robot vacuum cleaner, the electronic device 100 may generate a navigation map based on the spatial map, and perform cleaning along the moving path provided by the navigation map.

[0073] In an embodiment of the present disclosure, a spatial map can be generated by combining spatial scanning information or object information collected from different devices. The spatial map may further include information in the form of metadata indicating the device or position and / or viewing angle from which each piece of collected information is obtained. For example, in the case where different devices scan a space and detect objects from different positions and / or viewing angles, a spatial map can be generated by using images obtained by the corresponding devices or object recognition results obtained from the images. The spatial map can be marked or tagged for each object in the space with information indicating the object recognition result and the position and / or viewing angle from which the object recognition result is obtained.

[0074] Figure 5 is a graph used to describe object recognition results for images captured from different viewpoints.

[0075] Devices located in a space including at least one object may photograph (or capture) the object from different perspectives according to the position or operation method of each device and a camera provided in each device. Figure 5As shown, a first camera provided in the external device 300 may photograph an object in a space to obtain a first image captured from a first perspective. A second camera provided in the electronic device 100 may photograph an object in a space to obtain a second image captured from a second perspective.

[0076] The first camera provided in the external device 300, such as a smart home camera, can capture an object in a bird's eye view, a head-up view, or a high-angle view, but is stationary and therefore cannot capture the entire space. A general large-scale dataset for training a neural network model for performing a visual task mainly includes images captured in a bird's eye view, a head-up view, or a high-angle view, and can enable the neural network model to better recognize objects in the image than when trained based on a dataset including images captured in a low-angle view.

[0077] In contrast, an electronic device 100 such as a robot vacuum cleaner is movable, and therefore, a second camera provided in the electronic device 100 is advantageous in capturing various areas in a space, but is installed at a low height, and therefore, photographs an object at a low angle view. An image captured in a low angle view has a limited field of view, and therefore may contain only a portion of an object or may contain an object in a form that makes visual recognition difficult. Since it may be difficult to obtain a training data set for images captured in a low angle view, it is difficult to train a neural network model for performing visual tasks by using images captured in a low angle view as input, which inevitably reduces the accuracy and stability of object recognition. There may be practical difficulties in obtaining a training data set to supplement this problem because it is costly. It is difficult for an electronic device 100 such as a robot vacuum cleaner to receive high-resolution images for real-time calculations. In addition, when a high-performance camera is installed therein, cost issues arise.

[0078] Compared with the neural network model equipped in the electronic device 100 (e.g., a robot vacuum cleaner), the neural network model equipped in the external device 300 (e.g., a smart home camera) has relatively stable and accurate visual task performance. Hereinafter, a method for training a neural network model by using an object recognition result predicted by a neural network model equipped in the external device 300, which uses an image captured in a view (e.g., a low-angle view) in which the training data set is limited, will be described. In some embodiments, the method is used to train a neural network model by enhancing images representing objects captured by multiple cameras. One advantage of the method is to improve the accuracy of capturing images of objects in, for example, a background or space (e.g., a room). As data accumulates and the neural network model is learned using the accumulated data, the performance of the electronic device (e.g., a robot) (such as detection, tracking, and gesture recognition) is improved. The method of the present disclosure can be used in various applications, such as tracking members, recognizing gestures / actions (e.g., pointing gestures, cleaning specific locations), and detecting abnormal situations.

[0079] Figure 6 is a diagram for describing a process of generating training data by training a second neural network model using a first object recognition result predicted by a first neural network model using a first image as an input. Figure 6 The electronic device 100 is shown to train the second neural network model, but the present disclosure is not limited thereto, and the cloud server 200 that manages the second neural network model may train the second neural network model. Figures 6 to 12 In the description provided, as an example embodiment, the electronic device 100 trains a second neural network model.

[0080] Reference Figure 6 , the external device 300 includes a first camera configured to photograph a space including at least one object from a first perspective. The first neural network model equipped in the external device 300 can receive a first image captured by the first camera as an input and output a first object recognition result. The first image captured from the first perspective may be an image captured in a bird's-eye view, a head-up view, or a high-angle view. The electronic device 100 includes a second camera configured to photograph a space including at least one object from a second perspective. The second neural network model equipped in the electronic device 100 can receive a second image captured by the second camera as an input and output a second object recognition result. The second image captured from the second perspective may be an image captured in a low-angle view.

[0081] Figure 7 is a diagram for describing an example of a first neural network model or a second neural network model.

[0082] According to an embodiment of the present disclosure, the first neural network model or the second neural network model may be used to perform the following steps: Figure 7 A multi-task model for visual tasks as shown. The multi-task model may be a neural network including multiple layers. The multi-task model may include a shared backbone layer and additional layers for corresponding tasks, and the additional layers respectively receive the output of the shared backbone layer as input. The multi-task model may include a feature extractor for extracting features required for prediction, and multiple predictors for performing prediction by using the features extracted by the feature extractor. For example, for a visual model for performing tasks in the visual field, a multi-task model in which the prediction head of each task is connected to the feature extractor may be used. However, the type of multi-task model according to an embodiment of the present disclosure is not limited to Figure 7 Model shown.

[0083] The first neural network model provided in the external device 300 and the second neural network model provided in the electronic device 100 may be as follows: Figure 7 The multi-tasking model shown is, but not limited to, this.

[0084] refer to Figure 7 The multi-task model shown can input the features extracted by the shared backbone into the additional layer of the corresponding task to obtain the output value of the corresponding task. The prediction head for the corresponding task includes a detection and two-dimensional pose regression branch, a depth regression branch, and a re-identification (Re-ID) branch. Therefore, the output value of the prediction head for the corresponding task may include object detection information, two-dimensional pose information, three-dimensional pose information, re-identification feature information, etc. In this case, object tracking can be performed on continuous images through object detection information and re-identification feature information, and the recognition information of the object can be obtained by checking whether the identity of the object is maintained. Therefore, Figure 7 The multi-task model shown can receive an image as input and output an object recognition result for the image.

[0085] Depend on Figure 7 The object recognition result predicted by the multi-task model shown may include detection information of the object, two-dimensional pose information including image coordinates of feature points representing the position of the object in the image, three-dimensional pose information including spatial coordinates of feature points representing the position and orientation of the object, and recognition information of the object.

[0086] Return to reference Figure 6, a first image captured from a first perspective includes the entire object in space. In contrast, a second image captured from a second perspective includes only a portion of the back side of the same object in the same space. When a neural network model for performing a visual task performs a visual task on an image (e.g., the second image) that includes only a portion of an object, the neural network model may make errors such as failure to detect, identify, or locate the object. In the case where the neural network model for performing a visual task is based on self-supervised learning, it may be difficult to obtain a training data set due to object recognition errors, resulting in a vicious cycle in which the accuracy of the neural network model is not improved.

[0087] According to an embodiment of the present disclosure, in a case where an external device 300 and an electronic device 100 photograph the same object in the same space to obtain a first image from a first perspective and a second image from a second perspective, respectively, a first object recognition result output by a first neural network model equipped in the external device 300 for the first image can be converted by using the relationship between the first perspective and the second perspective, and the first object recognition result converted relative to the second perspective can be used as a label of the second image captured by the electronic device 100. A label is information related to raw data and refers to metadata that provides additional information about the raw data. Labeling is inputting or adding a label to the raw data, and according to the labeling method, annotations such as bounding boxes, polygon segmentation, points, or key points can be performed. For example, the electronic device 100 can obtain a label of the second image by converting the first object recognition result predicted by the first neural network model using the first image as input based on the conversion relationship between the first camera coordinate system of the external device 300 and the second camera coordinate system of the electronic device 100.

[0088] Figure 8 is a diagram for describing a process of converting a first object recognition result predicted by a first neural network model using a first image as input to obtain the first object recognition result converted relative to a second perspective. Figure 8 The conversion process can be performed by the electronic device 100 or the cloud server 200.

[0089] The first object recognition result may include detection information of the object, two-dimensional posture information in a first image coordinate system including image coordinates of feature points representing the position of the object in the first image, three-dimensional posture information in a first camera coordinate system including spatial coordinates of feature points representing the position and orientation of the object, and recognition information of the object.

[0090] Figure 8 The present invention shows a process of converting three-dimensional pose information in a first camera coordinate system included in a first object recognition result into two-dimensional pose information in a second image coordinate system, and converting the two-dimensional pose information into a bounding box (bbox).

[0091] First, based on the conversion relationship between the first camera coordinate system and the world coordinate system, the three-dimensional pose information in the first camera coordinate system is converted into the three-dimensional pose information in the world coordinate system. The conversion relationship between the first camera coordinate system and the world coordinate system can be identified by performing a first camera calibration to obtain the internal parameters and external parameters of the first camera. The internal parameters can be identified by the specifications of the camera (such as the focal length, principal point and skew coefficient of the camera). The external parameters involve the geometric relationship between the camera and the external space, such as the installation height and direction of the camera (for example, panning or tilting), and can vary according to the position and direction of the camera and how the world coordinate system is defined. The external parameters of the first camera can be identified by identifying the internal parameters of the first camera and then using the previously known spatial coordinates in the world coordinate system and the matching pairs of image coordinates in the image coordinate system of the first camera corresponding to the spatial coordinates.

[0092] For example, the external parameters of the first camera installed on the external device 300 (such as a smart home camera) can be identified by using the internal parameters of the first camera and the spatial coordinates of at least four feature points in the world coordinate system and the image coordinates in the image coordinate system of the first camera corresponding to the feature points respectively. In this case, the matching pair can be obtained by detecting a plurality of feature points within the field of view of the first camera of the external device 300 using the mobile electronic device 100. For example, in the case where the electronic device 100 is a robot vacuum cleaner, the center point of a certain detection area can be used as a feature point, or a mark attached to the robot vacuum cleaner or an external feature of the robot vacuum cleaner can be used as a feature point. The spatial coordinates of the feature points in the world coordinate system can be obtained by the height of the robot vacuum cleaner and the position of the feature points within the robot vacuum cleaner and simultaneous localization and mapping (SLAM), and the image coordinates of the feature points in the image coordinate system of the first camera can be obtained by the two-dimensional coordinates of the feature points detected in the smart home camera image.

[0093] Based on the conversion relationship between the world coordinate system and the second camera coordinate system, the three-dimensional pose information in the world coordinate system is converted into the three-dimensional pose information in the second camera coordinate system. The conversion relationship between the world coordinate system and the second camera coordinate system can be identified by performing a second camera calibration to obtain external parameters of the second camera, and can be represented by a rotation matrix and a translation matrix for converting the world coordinate system into the second camera coordinate system.

[0094] By projecting the spatial coordinates in the second camera coordinate system to the second image coordinate system of the second camera, the three-dimensional pose information in the second camera coordinate system is converted into two-dimensional pose information in the second image coordinate system. The projection matrix can be identified by performing second camera calibration to obtain the internal parameters of the second camera.

[0095] By allowing a preset margin value for each minimum value and maximum value of x and y of the two-dimensional pose information in the second image coordinate system, bounding box (bbox) information for estimating an object area may be generated from the two-dimensional pose information in the second image coordinate system.

[0096] Therefore, a first object recognition result transformed relative to the second viewing angle can be obtained, the first object recognition result including detection information and recognition information of the object, two-dimensional posture information in the second image coordinate system, three-dimensional posture information in the second camera coordinate system, and bounding box information capable of identifying the existence of the object in the second image. Return to reference Figure 6 , the electronic device 100 may generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result converted relative to the second viewing angle. The first image and the second image corresponding to each other means that the relationship between them is most suitable for generating training data for the second image by using the object recognition result predicted by the neural network model using the first image as input. The electronic device 100 can identify a pair of timestamps corresponding to each other by comparing the timestamp corresponding to the first object recognition result with the timestamp of the second image. The timestamp corresponding to the first object recognition result may be the timestamp of the first image. When it is determined that the object exists in the second image corresponding to the first image, labeling may be performed on the second image based on the first object recognition result converted relative to the second viewing angle.

[0097] For example, in the case where the electronic device 100 for performing a task by recognizing the posture information of an object is a robot vacuum cleaner, it is necessary to train the neural network model equipped in the robot vacuum cleaner to better recognize the posture information of the object, but there are limitations on the training data set for training the neural network model equipped in the robot vacuum cleaner. To this end, the electronic device 100 can generate an image captured at a low angle view as training data by using the object recognition result predicted by the first neural network model equipped in the external device 300.

[0098] As referenced above Figure 8As described, the electronic device 100 can obtain the first object recognition result converted relative to the second perspective by converting the first object recognition result predicted by the first neural network model using the first image as input. For example, the electronic device 100 can convert the three-dimensional posture information including the spatial coordinates of the feature points representing the position and direction of the object recognized in the first image into two-dimensional posture information including the image coordinates of the feature points representing the position of the object recognized in the second image based on the conversion relationship between the first camera coordinate system corresponding to the first camera configured to capture the image from the first perspective and the second camera coordinate system corresponding to the second camera configured to capture the image from the second perspective. The electronic device 100 can generate training data by performing labeling on the second image, so that the two-dimensional posture information included in the first object recognition result converted relative to the second perspective is used as the two-dimensional posture information of the object in the second image.

[0099] The electronic device 100 can generate training data by performing labeling on the second image so that, among the two-dimensional posture information included in the first object recognition result converted relative to the second viewing angle, the first part corresponding to the second image is distinguished from the second part that does not correspond to the second image. In the case where a part of the two-dimensional posture information included in the first object recognition result converted relative to the second viewing angle is not visible in the second image, the second image can be labeled with the invisible part and the visible part in the second image. For example, in the case where the two-dimensional posture information is skeleton information about the entire body of the object and even the upper body of the object is not visible in the second image, the electronic device 100 can perform labeling on the second image based on the two-dimensional posture information, which is skeleton information about the entire body of the object. In this case, the electronic device 100 can generate training data by performing labeling so that for (x, y, visibility) representing the position and visibility value of each joint point constituting the skeleton information, the visibility value corresponding to each joint point of the lower body visible in the second image is set to "1", and the visibility value corresponding to each joint point of the upper body not visible in the second image is set to "0".

[0100] According to an embodiment of the present disclosure, the electronic device 100 may generate a plurality of second images as training data. Figures 9 to 11 Provide a description.

[0101] Fig. 9 is a diagram for describing an example of generating training data for each of a plurality of second images captured from a second perspective.

[0102] Fig. 9The second camera of the electronic device 100 is shown to obtain a plurality of second images captured from a second perspective by capturing a dynamic object at a preset time interval while maintaining the second perspective. In the case where the plurality of second images are captured by the second camera from the second perspective as described above, the electronic device 100 can generate training data by performing labeling on each of the plurality of second images based on the first object recognition result converted relative to the second perspective.

[0103] like Fig. 9 As shown, when the second camera captures images for a specific time period while maintaining a second viewing angle, when t=1, an image in which the object appears at the right end of the field of view of the electronic device 100 can be captured, when t=2, an image in which the object is located at the center of the field of view of the electronic device 100 can be captured, and when t=3, an image in which the object disappears from the left end of the field of view of the electronic device 100 can be captured. The second camera of the electronic device 100 captures the second image at a low-angle view, mainly focusing on the lower half of the object rather than the entire body, and it may be difficult to detect and recognize the posture information of the object. In this case, by using a first object recognition result predicted by a first neural network model using a first image obtained by a first camera of an external device 300 that captures the object as an input, labeling based on the first object recognition result converted relative to the second viewing angle can be performed on each of the multiple second images. Referring to Fig. 9 , for t=1 and t=3, only a portion of the two-dimensional pose information included in the first object recognition result converted relative to the second viewing angle can be used as a label for the pose information of the object in the second image captured from the second viewing angle. For t=2, joint point information corresponding to parts of the lower body and the upper body in the two-dimensional pose information included in the first object recognition result converted relative to the second viewing angle can be used as a label for the pose information of the object in the second image captured from the second viewing angle.

[0104] Fig.10 and Fig.11 is a diagram for describing an example of generating training data for each of a plurality of second images captured from a plurality of different second perspectives, respectively.

[0105] Fig.10 and Fig.11 Each shows that the second camera of the electronic device 100 obtains a plurality of second images by capturing an object from a plurality of different second viewing angles. Fig.10 1 shows a plurality of second images respectively captured from a plurality of different second viewing angles by a second camera mounted on the electronic device 100 while rotating around a static object. Fig.10 As shown, the second camera of the electronic device 100 captures a second image of mainly the lower half of the subject rather than the entire body by capturing the back, left, and right sides of the subject at a low angle view, and it may be difficult to detect and recognize posture information of the subject. Fig.11 2 shows a plurality of second images captured by the second camera of the electronic device 100 from a plurality of different second viewing angles respectively while changing the distance to the static object. Fig.11 As shown, the second camera of the electronic device 100 captures a second image of mainly the lower half of the subject rather than the entire body by capturing the subject at a low angle view while getting closer or farther away from the subject, so that the subject appears larger or smaller in the image and it may be difficult to detect and recognize the subject's posture information.

[0106] In this way, when the second camera captures a plurality of second images from a plurality of different second viewing angles, respectively, the electronic device 100 can transform the first object recognition result predicted by the first neural network model of the external device 300 based on the transformation relationship between the first camera coordinate system and the second camera coordinate system for each of the plurality of second viewing angles. This is because the first object recognition result transformed based on each of the plurality of different second viewing angles is required for labeling. The electronic device 100 can generate training data by performing labeling on the plurality of second images based on the first object recognition results transformed relative to the plurality of second viewing angles, respectively.

[0107] In an embodiment, the electronic device 100 may generate training data at any time or according to the user's settings. In addition, when the electronic device 100 equipped with the second camera recognizes an object at a level below a preset standard, or when the electronic device 100 moves to a position specified in a spatial map used by the electronic device 100, the electronic device 100 may generate training data. For example, when the electronic device 100 is at a position specified in the spatial map, the field of view of the first camera of the external device 300 and the field of view of the second camera of the electronic device 100 may overlap with each other. In addition, the position specified in the spatial map may be an area where an object frequently moves or a position arbitrarily marked by a user.

[0108] Return to reference Figure 6 , the electronic device 100 can train the second neural network model by using the generated training data.

[0109] Fig.12 is a diagram for describing a process of training a second neural network model by using training data generated by using an output of a first neural network model.

[0110] As described above, a first object recognition result predicted by a first neural network model of the external device 300 is obtained from a first image captured at a first perspective, and therefore, the first object recognition result can be converted relative to a second perspective, and labeling can be performed using the first object recognition result converted relative to the second perspective to generate training data for a second neural network model of the electronic device 100.

[0111] The electronic device 100 can train the second neural network model by using the generated training data. In this case, the electronic device 100 can fine-tune the prediction head for each task of the second neural network model. Fine-tuning refers to updating the existing trained neural network model to adapt to the purpose by modifying the architecture of the existing trained neural network model. The second neural network model trained in this way can predict a second object recognition result with improved object recognition accuracy and stability.

[0112] Fig.13 is a flowchart for describing a method for training a neural network model.

[0113] The electronic device 100 or the cloud server 200 may obtain a first object recognition result predicted by a first neural network model using a first image captured by a first camera as an input, the first camera capturing a space including at least one object from a first perspective (S1310).

[0114] According to an embodiment of the present disclosure, a first camera provided in the external device 300 may photograph a space including at least one object from a first perspective. A first neural network model provided in the external device 300 may predict and output a first object recognition result by using a first image captured from the first perspective as an input. The electronic device 100 or the cloud server 200 may receive the first object recognition result from the external device 300.

[0115] According to an embodiment of the present disclosure, a first camera provided in an external device 300 may photograph a space including at least one object from a first perspective. The external device 300 may send a first image captured by the first camera from the first perspective to a cloud server 200 that manages a first neural network model. The cloud server 200 may receive the first image captured from the first perspective from the external device 300. The first neural network model managed by the cloud server 200 may predict and output a first object recognition result by using the first image captured from the first perspective as an input. The cloud server 200 may obtain a first object recognition result predicted by the first neural network model.

[0116] The electronic device 100 or the cloud server 200 may convert the obtained first object recognition result based on the conversion relationship between the first camera coordinate system corresponding to the first camera and the second camera coordinate system corresponding to the second camera capturing the space from the second perspective (S1320). Therefore, the first recognition result based on the first perspective may be converted relative to the second perspective. The electronic device 100 or the cloud server 200 may use the first object recognition result converted relative to the second perspective as a label of the second image.

[0117] The first object recognition result may include detection information of the object, two-dimensional posture information in a first image coordinate system including image coordinates of feature points representing the position of the object in the first image, three-dimensional posture information in a first camera coordinate system including spatial coordinates of feature points representing the position and orientation of the object, and recognition information of the object.

[0118] The electronic device 100 or the cloud server 200 converts the three-dimensional pose information in the first camera coordinate system into the three-dimensional pose information in the world coordinate system based on the conversion relationship between the first camera coordinate system and the world coordinate system. The conversion relationship between the first camera coordinate system and the world coordinate system can be obtained by performing a first camera calibration to obtain the internal parameters of the first camera and the external parameters of the first camera using a matching pair of the spatial coordinates of a certain number of feature points in the world coordinate system and the image coordinates of the feature points in the image coordinate system of the first camera. Here, the feature points can be detected by moving the electronic device 100 within the field of view of the first camera and detecting the center point of a specific detection area, a mark attached to the electronic device 100, or a feature of the appearance of the electronic device 100.

[0119] The electronic device 100 or the cloud server 200 converts the 3D pose information in the world coordinate system into the 3D pose information in the second camera coordinate system based on the conversion relationship between the world coordinate system and the second camera coordinate system. The electronic device 100 or the cloud server 200 converts the 3D pose information in the second camera coordinate system into the 2D pose information in the second image coordinate system by projecting the spatial coordinates in the second camera coordinate system into the second image coordinate system of the second camera. The electronic device 100 or the cloud server 200 generates bounding box (bbox) information from the 2D pose information in the second image coordinate system.

[0120] Therefore, a first object recognition result transformed relative to the second viewing angle can be obtained, and the first object recognition result includes detection information and recognition information of the object, two-dimensional posture information in the second image coordinate system, three-dimensional posture information in the second camera coordinate system, and bounding box information capable of identifying the existence of the object in the second image. In the first object recognition result transformed relative to the second viewing angle, the bounding box information can be used as a label for estimating the object area in the second image. The posture information can be used as a label for identifying the posture of the object in the second image. The recognition information of the object can be used as a label for identifying whether the object exists in a continuous second image.

[0121] The electronic device 100 or the cloud server 200 may generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result converted relative to the second viewing angle, the second image being captured by the second camera (S1330). The electronic device 100 or the cloud server 200 may identify a pair of timestamps corresponding to each other by comparing the timestamp corresponding to the first object recognition result with the timestamp of the second image. The timestamp corresponding to the first object recognition result may be the timestamp of the first image. When it is determined that the object exists in the second image corresponding to the first image, the electronic device 100 or the cloud server 200 may perform labeling on the second image based on the first object recognition result converted relative to the second viewing angle, wherein the second viewing angle is obtained by converting the first object recognition result. The electronic device 100 may generate training data by performing labeling on the second image so that, among the two-dimensional posture information included in the first object recognition result converted relative to the second viewing angle, the first portion corresponding to the second image is distinguished from the second portion not corresponding to the second image.

[0122] For example, in the case where the electronic device 100 is a robot vacuum cleaner for recognizing the posture information of the user and performing cleaning, in order to train the neural network model equipped in the electronic device 100 to better recognize the posture information of the object, the electronic device 100 can generate an image captured at a low angle view as training data by using the object recognition result predicted by the neural network model equipped in the external device 300. The electronic device 100 can convert the three-dimensional posture information including the spatial coordinates of the feature points representing the position and direction of the object recognized in the first image into the two-dimensional posture information including the image coordinates of the feature points representing the position of the object recognized in the second image based on the conversion relationship between the first camera coordinate system corresponding to the first camera configured to capture the image from the first perspective and the second camera coordinate system corresponding to the second camera configured to capture the image from the second perspective. The electronic device 100 can generate the training data by performing labeling on the second image so that the two-dimensional posture information included in the first object recognition result converted relative to the second perspective is used as the two-dimensional posture information of the object in the second image.

[0123] According to an embodiment of the present disclosure, in the case where the second camera captures a plurality of second images from a second viewing angle, the electronic device 100 may generate training data by performing labeling on each of the plurality of second images based on the first object recognition result converted with respect to the second viewing angle.

[0124] According to an embodiment of the present disclosure, when the second camera captures a plurality of second images from a plurality of different second viewing angles, the electronic device 100 may transform the predicted first object recognition result based on the transformation relationship between the first camera coordinate system and the second camera coordinate system for each of the plurality of second viewing angles. The electronic device 100 may generate training data by performing labeling on the plurality of second images, respectively, based on the first object recognition results transformed relative to the plurality of second viewing angles.

[0125] According to an embodiment of the present disclosure, when the electronic device 100 equipped with the second camera recognizes an object at a level lower than a preset standard, or when the electronic device 100 moves to a position specified in a spatial map used by the electronic device 100, the electronic device 100 may generate training data. For example, when the electronic device 100 is at a position specified in the spatial map, the field of view of the first camera of the external device 300 and the field of view of the second camera of the electronic device 100 may overlap with each other. In addition, the position specified in the spatial map may be an area where an object frequently moves or a position arbitrarily marked by a user.

[0126] The electronic device 100 or the cloud server 200 may train the second neural network model by using the generated training data (S1340). The electronic device 100 or the cloud server 200 may fine-tune the prediction head of each task of the second neural network model by using the generated training data.

[0127] Fig.14 is a diagram illustrating an example of training a neural network model on the electronic device 100.

[0128] refer to Fig.14 , the external device 300 is equipped with a first neural network model, and the electronic device 100 is equipped with a second neural network model. The first camera provided in the external device 300 may photograph a space including at least one object from a first perspective. The first neural network model provided in the external device 300 may predict and output a first object recognition result by using a first image captured from the first perspective as an input. The external device 300 may send the first object recognition result predicted by the first neural network model to the electronic device 100. The electronic device 100 may receive the first object recognition result from the external device 300.

[0129] The electronic device 100 may convert the first object recognition result received from the external device 300 based on the conversion relationship between the second camera coordinate system corresponding to the second camera capturing the same space from the second viewing angle and the first camera coordinate system corresponding to the first camera of the external device 300. Therefore, the first recognition result received from the external device 300 may be converted relative to the second viewing angle. The electronic device 100 may use the first object recognition result converted relative to the second viewing angle as a label of the second image.

[0130] The electronic device 100 may generate training data by performing labeling on a second image corresponding to a first image captured from a first perspective by a first camera of the external device 300 based on the first object recognition result converted with respect to the second perspective. The electronic device 100 may train a second neural network model by using the generated training data.

[0131] Fig.15 is a diagram showing an example of training a neural network model on the cloud server 200.

[0132] Reference Fig.15 , the external device 300 is equipped with a first neural network model, and the cloud server 200 and the electronic device 100 are equipped with a second neural network model. The first camera set in the external device 300 can shoot a space including at least one object from a first perspective. The first neural network model equipped in the external device 300 can predict and output a first object recognition result by using a first image captured from the first perspective as an input. The external device 300 can send the first object recognition result predicted by the first neural network model to the cloud server 200. The cloud server 200 can receive the first object recognition result from the external device 300.

[0133] The second camera provided in the electronic device 100 may obtain a second image by capturing the same space photographed by the first camera provided in the external device 100 from a second perspective. The electronic device 100 may transmit the second image to the cloud server 200. The cloud server 200 may receive the second image from the electronic device 100.

[0134] The cloud server 200 may convert the first object recognition result received from the external device 300 based on the conversion relationship between the first camera coordinate system corresponding to the first camera of the external device 300 and the second camera coordinate system corresponding to the second camera of the electronic device 100. To this end, the cloud server 200 may store information about the internal parameters and external parameters of the first camera related to the first camera coordinate system and information about the internal parameters and external parameters of the second camera related to the second camera coordinate system. The cloud server 200 may use the first object recognition result converted with respect to the second viewing angle as a label of the second image.

[0135] The cloud server 200 may generate training data by performing labeling on a second image corresponding to a first image captured from a first perspective by a first camera of the external device 300 based on the first object recognition result converted with respect to the second perspective. The cloud server 200 may transmit the generated training data to the electronic device 100 equipped with the second neural network model. In this case, the electronic device 100 may train the second neural network model equipped therein by using the received training data.

[0136] The cloud server 200 may train the second neural network model by using the generated training data. When the second neural network model is trained by the cloud server 200, the cloud server 200 may send updated network parameter values ​​of the trained second neural network model to the electronic device 100 equipped with the second neural network model. The electronic device 100 may update the second neural network model based on the received updated network parameter values.

[0137] Fig.16 is a diagram for illustrating an example of training a neural network model on the cloud server 200.

[0138] refer to Fig.16 , the cloud server 200 is equipped with a first neural network model and a second neural network model, and the electronic device 100 is equipped with a second neural network model. A first camera provided in the external device 300 may photograph a space including at least one object from a first perspective. The external device 300 may transmit a first image captured from a first perspective to the cloud server 200. The cloud server 200 may receive the first image from the external device 300.

[0139] The cloud server 200 may input the first image into the first neural network model to obtain a first object recognition result predicted by the first neural network model.

[0140] The second camera provided in the electronic device 100 may obtain a second image by capturing the same space photographed by the first camera provided in the external device 100 from a second perspective. The electronic device 100 may transmit the second image to the cloud server 200. The cloud server 200 may receive the second image from the electronic device 100.

[0141] The cloud server 200 may transform the first object recognition result predicted by the first neural network model based on a transformation relationship between a first camera coordinate system corresponding to the first camera of the external device 300 and a second camera coordinate system corresponding to the second camera of the electronic device 100. The cloud server 200 may use the first object recognition result transformed relative to the second viewing angle as a label of the second image.

[0142] The cloud server 200 may generate training data by performing labeling on a second image corresponding to a first image captured from a first perspective by a first camera of the external device 300 based on the first object recognition result converted with respect to the second perspective. The cloud server 200 may transmit the generated training data to the electronic device 100 equipped with the second neural network model. In this case, the electronic device 100 may train the second neural network model equipped therein by using the received training data.

[0143] The cloud server 200 may train the second neural network model by using the generated training data. When the second neural network model is trained by the cloud server 200, the cloud server 200 may send updated network parameter values ​​of the trained second neural network model to the electronic device 100 equipped with the second neural network model. The electronic device 100 may update the second neural network model based on the received updated network parameter values.

[0144] Fig.17 is a diagram for describing an example in which the electronic device 100 performs a task by recognizing a user's gesture.

[0145] refer to Fig.17 , the electronic device 100 is a robot vacuum cleaner, which is configured to recognize the posture information of the user and perform cleaning. When the neural network model equipped in the robot vacuum cleaner is trained to recognize the posture information of the user according to the above method, the robot vacuum cleaner can recognize the posture information or gesture information of the user even through an image captured at a low angle viewing angle.

[0146] like Fig.17 As shown, when a user gives a voice command saying "Hi, cleaning robot! Please clean the room there", the robot vacuum cleaner is able to recognize that the cleaning command has been initiated by the user through a neural network model configured to perform natural language processing, but may not be able to specify the specific location of "the room there". In this case, the robot vacuum cleaner can recognize the room the user is pointing to by checking the user's gesture information estimated by a neural network model configured to perform object detection and gesture recognition. As a result, the robot vacuum cleaner can confirm the command for cleaning the room identified according to the user's estimated gesture information and then perform cleaning.

[0147] Fig.18 and Fig.19 is a block diagram showing a configuration of an electronic device 100 according to an embodiment of the present disclosure.

[0148] refer to Fig.18The electronic device 100 according to an embodiment of the present disclosure may include a memory 110, a processor 120, a camera 131, and a communication unit 140, but is not limited thereto and may further include general components. Fig.19 As shown, in addition to the memory 110, the processor 120, the camera 131 and the communication unit 140, the electronic device 100 may further include a sensing unit 130, an input / output unit 150 and a driving unit 160, and the sensing unit 130 includes the camera 131. Fig.18 and Fig.19 Describe each component in detail.

[0149] The memory 110 according to an embodiment of the present disclosure may store a program for the processor 120 to perform processing and control, and may store data input to or generated by the electronic device 100. The memory 110 may store instructions, data structures, and program codes that may be read by the processor 120. In an embodiment of the present disclosure, the operations performed by the processor 120 may be implemented by executing instructions or codes of the program stored in the memory 110.

[0150] The memory 110 according to an embodiment of the present disclosure may include a flash memory, a hard disk memory, a multimedia card micro memory, a card-type memory (for example, an SD memory, an XD memory, etc.), a non-volatile memory including at least one of a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), a programmable ROM (PROM), a magnetic memory, a magnetic disk, or an optical disk, and a volatile memory such as a random access memory (RAM) or a static RAM (SRAM).

[0151] The memory 110 according to an embodiment of the present disclosure may store at least one instruction and / or program for controlling the electronic device 100 to train a neural network model. For example, the memory 110 may store an object recognition result conversion module, a training data generation module, a training module, and the like.

[0152] The processor 120 according to an embodiment of the present disclosure may execute instructions stored in the memory 110 or programmed software modules to control the operation or function of the electronic device 100 to perform tasks. The processor 120 may include hardware components that perform arithmetic operations, logical operations, input / output operations, and signal processing. The processor 120 may execute at least one instruction stored in the memory 110 to control the overall operation of the electronic device 100 to train a neural network model and perform tasks by using the trained neural network model. The processor 120 may execute a program stored in the memory 110 to control the camera 131 or the sensing unit 130, the communication unit 140, the input / output unit 150, and the driving unit 160 including the camera 131.

[0153] For example, the processor 120 according to an embodiment of the present disclosure may include, but is not limited to, a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), an application processor, a neural processing unit, or at least one of a dedicated artificial intelligence processor designed with a hardware structure dedicated to processing an artificial intelligence model. Each processor constituting the processor 120 may be a dedicated processor for performing a certain function.

[0154] The AI ​​processor according to an embodiment of the present disclosure may perform calculations and controls using an AI model to process tasks that the electronic device 100 is configured to perform. The AI ​​processor may be manufactured in the form of a dedicated hardware chip for AI, or may be manufactured as part of a general-purpose processor (e.g., a CPU or an application processor) or a dedicated graphics processor (e.g., a GPU) and mounted on the electronic device 100.

[0155] The sensing unit 130 according to an embodiment of the present disclosure may include a plurality of sensors configured to detect information about the surrounding environment of the electronic device 100. For example, the sensing unit 130 may include a camera 131, a light detection and ranging (LiDAR) sensor 132, an infrared sensor 133, an ultrasonic sensor 134, a ToF sensor 135, a gyro sensor 136, etc., but is not limited thereto.

[0156] The camera 131 according to an embodiment of the present disclosure may include a stereo camera, a monochrome camera, a wide-angle camera, a panoramic camera, a three-dimensional vision sensor, and the like.

[0157] The LiDAR sensor 132 can detect the distance and various physical characteristics of an object by emitting laser light toward the target. The LiDAR sensor 132 can be used to detect surrounding objects, geographical features, etc., and model them into a three-dimensional image.

[0158] The infrared sensor 133 may be any one of an active infrared sensor configured to emit infrared radiation and detect changes due to light blocking and a passive infrared sensor having no light emitter and configured to detect only changes in infrared rays received from the outside. For example, an infrared proximity sensor may be installed around the wheels of the electronic device 100 to function as an anti-fall sensor by emitting infrared rays toward the floor and then receiving them.

[0159] The ultrasonic sensor 134 may measure the distance to an object by using ultrasonic waves, and may emit and detect ultrasonic pulses that convey information about the proximity of an object. The ultrasonic sensor 134 may be used to detect nearby objects and transparent objects.

[0160] The ToF sensor 135 can calculate the distance that the light emitted toward the object is reflected and returned according to time to obtain the three-dimensional effect, movement and spatial information of the object. The ToF sensor 135 can realize advanced recognition of objects in complex spaces and dark places and even obstacles in front of the eyes to allow the electronic device 100 to avoid obstacles.

[0161] The gyro sensor 136 may detect an angular velocity and may be used to measure the position of the electronic device 100 and set the direction of the electronic device 100 .

[0162] According to an embodiment of the present disclosure, the sensing unit 130 may be used to photograph a space including at least one object and generate spatial information by using the camera 131. The electronic device 100 may obtain spatial information about a space including at least one object by obtaining spatial scanning information or object information using a plurality of sensors of the same type or different types among the camera 131, the LiDAR sensor 132, the infrared sensor 133, the ultrasonic sensor 134, the ToF sensor 135, and the gyro sensor 136. However, even in the case where the electronic device 100 includes a sensor configured to recognize depth information, such as the LiDAR sensor 132 or the ToF sensor 135, it may be difficult for the electronic device 100 to detect the object due to the direction of the object or the distance or relative position between the object and the electronic device 100.

[0163] The communication unit 140 may include one or more components configured to implement communication between the electronic device 100 and an external device such as the cloud server 200, the IoT devices 300-1, 300-2, and 300-3, or a user terminal. For example, the communication unit 140 may include a short-range wireless communication unit 141, a mobile communication unit 142, etc., but is not limited thereto.

[0164] The short-range wireless communication unit 141 may include a Bluetooth communication unit, a Bluetooth low energy (BLE) communication unit, a near field communication unit, a wireless local area network (WLAN) (e.g., Wi-Fi) communication unit, a Zigbee communication unit, an Ant+ communication unit, a Wi-Fi Direct (WFD) communication unit, an ultra-wideband (UWB) communication unit, an infrared data association (IrDA) communication unit, a microwave (uWave) communication unit, etc., but is not limited thereto.

[0165] The mobile communication unit 142 transmits and receives radio signals to and from at least one of a base station, an external terminal or a server over a mobile communication network. Here, the radio signal may include a voice call signal, a video call signal or various types of data transmitted and received according to text / multimedia messages.

[0166] like Fig.19 As shown, the electronic device 100 may further include an input / output unit 150 and a driving unit 160, and although Fig.19 Not shown, but components such as a power supply unit may also be included.

[0167] The input / output unit 150 may include an input unit 151 and an output unit 153. The input / output unit 150 may include the input unit 151 and the output unit 153 separated from each other, or may include a component in which they are integrated, such as a touch screen. The input / output unit 150 may receive input information from a user and provide output information to the user.

[0168] The input unit 151 may refer to a unit through which a user inputs data for controlling the electronic device 100. For example, the input unit 151 may include a keyboard, a touch panel (e.g., a touch-type capacitive touch panel, a pressure-type resistive overlay touch panel, an infrared sensor type touch panel, a surface acoustic wave conduction touch panel, an integrated tension measurement touch panel, a piezoelectric effect type touch panel), etc. In addition, the input unit 151 may include a jog wheel, a jog switch, etc., but is not limited thereto.

[0169] The output unit 153 may output an audio signal, a video signal or a vibration signal, and may include a display unit, an audio output unit and a vibration motor.

[0170] The display unit may display information processed by the electronic device 100. For example, the display unit may display a user interface for receiving user manipulation. In the case where the display unit and the touch panel constitute a layer structure to form a touch screen, the display may be used as an input device in addition to an output device. The display unit may include at least one of a liquid crystal display, a thin film transistor liquid crystal display, an organic light emitting diode, a flexible display, or a three-dimensional (3D) display. Depending on the embodiment of the electronic device 100, the electronic device 100 may include two or more display units.

[0171] The audio output unit may output audio data stored in the memory 110. The audio output unit may output an audio signal related to a function performed by the electronic device 100. The audio output unit may include a speaker, a buzzer, and the like.

[0172] The vibration motor may output a vibration signal. For example, the vibration motor may output a vibration signal corresponding to the output of audio data or video data. When a touch is input to the touch screen, the vibration motor may output a vibration signal.

[0173] The driving unit 160 may include components for operating (driving) the electronic device 100 and internal devices of the electronic device 100. In the case where the electronic device 100 is a robot vacuum cleaner, the driving unit 160 may include a suction unit, a travel unit, etc., but is not limited thereto, and the driving unit 160 may vary according to the type of the electronic device 100.

[0174] The suction unit plays a role of collecting dust on the floor while sucking air, and may include, but is not limited to, a rotating brush or broom, a rotating brush motor, an air suction port, a filter, a dust collection chamber, an air exhaust port, etc. The suction unit may be installed in a structure in which an additional brush for sweeping dust from corners is rotatable.

[0175] The traveling unit may include, but is not limited to, a motor that rotates each wheel installed in the electronic device 100 , and a timing belt installed to transfer power generated from the wheel.

[0176] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to obtain, through the communication unit 140, a first object recognition result predicted by the first neural network model using a first image captured from a first perspective by a camera of an external device 300 capturing a space including at least one object as an input. The processor 120 may execute at least one instruction stored in the memory 110 to transform the predicted first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the camera of the external device 300 and a second camera coordinate system corresponding to the camera 131 capturing the space from a second perspective. The processor 120 may execute at least one instruction stored in the memory 110 to generate training data by performing labeling on a second image corresponding to the first image and captured by the camera 131 based on the first object recognition result transformed with respect to the second perspective. The processor 120 may execute at least one instruction stored in the memory 110 to train the second neural network model by using the generated training data.

[0177] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to convert the three-dimensional pose information including the spatial coordinates of the feature points indicating the position and direction of the object recognized in the first image into the two-dimensional pose information including the image coordinates of the feature points indicating the position of the object recognized in the second image based on the conversion relationship between the first camera coordinate system corresponding to the camera of the external device 300 and the second camera coordinate system corresponding to the camera 131. The processor 120 may execute at least one instruction stored in the memory 110 to generate training data by performing labeling on the second image, so that the two-dimensional pose information included in the first object recognition result converted relative to the second viewing angle is used as the two-dimensional pose information of the object in the second image.

[0178] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to convert the three-dimensional pose information in the first camera coordinate system into the three-dimensional pose information in the world coordinate system based on the conversion relationship between the first camera coordinate system and the world coordinate system. The processor 120 may execute at least one instruction stored in the memory 110 to convert the three-dimensional pose information in the world coordinate system into the three-dimensional pose information in the second camera coordinate system based on the conversion relationship between the world coordinate system and the second camera coordinate system. The processor 120 may execute at least one instruction stored in the memory 110 to convert the three-dimensional pose information in the second camera coordinate system into the two-dimensional pose information in the second image coordinate system by projecting the spatial coordinates in the second camera coordinate system into the second image coordinate system of the second camera.

[0179] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to obtain a conversion relationship between the first camera coordinate system and the world coordinate system by performing a first camera calibration, which obtains the internal parameters of the camera of the external device 300 and the external parameters of the camera of the external device 300 using matching pairs of spatial coordinates of a certain number of feature points in the world coordinate system and image coordinates of the feature points in the image coordinate system of the camera of the external device 300. Here, the feature points may be detected by moving the electronic device 100 within the field of view of the camera of the external device 300 and detecting a center point of a specific detection area, a marker attached to the electronic device 100, or a feature of the appearance of the electronic device 100.

[0180] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to generate training data by performing labeling on the second image, so that in the two-dimensional posture information included in the first object recognition result converted relative to the second perspective, the first part corresponding to the second image and the second part not corresponding to the second image are distinguished.

[0181] According to an embodiment of the present disclosure, the processor 120 can execute at least one instruction stored in the memory 110 to generate training data by labeling each of the multiple second images based on the first object recognition result converted relative to the second perspective when the camera 131 captures multiple second images from a second perspective.

[0182] According to an embodiment of the present disclosure, the processor 120 may execute at least one instruction stored in the memory 110 to transform the predicted first object recognition result based on the transformation relationship between the first camera coordinate system and the second camera coordinate system for each of the plurality of second perspectives when the camera 131 captures the plurality of second images respectively from the plurality of different second perspectives. The processor 120 may execute at least one instruction stored in the memory 110 to generate training data by performing labeling on each of the plurality of second images based on the first object recognition result transformed relative to each of the plurality of second perspectives.

[0183] According to an embodiment of the present disclosure, when the electronic device 100 recognizes an object below a preset standard or moves to a position specified in a spatial map used by the electronic device 100, the processor 120 may execute at least one instruction stored in the memory 110 to generate training data.

[0184] Fig. 20 is a block diagram showing a configuration of the server 200 according to an embodiment of the present disclosure.

[0185] The operation of the above-mentioned electronic device 100 can be performed by the server 200 in a similar manner. The server 200 according to an embodiment of the present disclosure may include a memory 210, a processor 220, a communication unit 230, and a storage device 240. The components of the server 200 may correspond to Fig.18 and Fig.19 The memory 110, the processor 120, and the communication unit 140 of the electronic device 100, and thus, the description provided above will be omitted.

[0186] The memory 210 may store various data, programs or applications for operating and controlling the server 200. At least one instruction or application stored in the memory 210 may be executed by the processor 220. The memory 210 may store modules configured to perform the same functions as the modules stored in the electronic device 100. For example, the memory 210 may store an object recognition result conversion module, a training data generation module, a training module, a prediction module, and data and program instruction codes corresponding thereto. The processor 220 may control the overall operation of the server 200. The processor 220 according to an embodiment of the present disclosure may execute at least one instruction stored in the memory 210.

[0187] The communication unit 230 may include one or more components configured to enable communication through a LAN, a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof.

[0188] The storage device 240 may store the first neural network model or the second neural network model. The storage device 240 may store a training data set for training various AI models.

[0189] The server 200 according to an embodiment of the present disclosure may be a device having higher computing performance than that of the electronic device 100 and thus capable of performing a larger amount of computing. The server 200 may perform training of an AI model requiring a relatively large amount of computing.

[0190] According to an embodiment of the present disclosure, the processor 220 may execute at least one instruction stored in the memory 210 to obtain a first object recognition result predicted by a first neural network model using a first image captured by a first camera as an input, the first camera capturing a space including at least one object from a first perspective. The processor 220 may execute at least one instruction stored in the memory 210 to obtain, through the communication unit 230, a second image captured by a second camera capturing a space from a second perspective. The processor 220 may execute at least one instruction stored in the memory 210 to transform the predicted first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to the second camera. The processor 220 may execute at least one instruction stored in the memory 210 to generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective. The processor 220 may execute at least one instruction stored in the memory 210 to train the second neural network model by using the generated training data.

[0191] According to an embodiment of the present disclosure, the processor 220 may execute at least one instruction stored in the memory 210 to transmit the network parameter value of the trained second neural network model to the electronic device 100 equipped with the second neural network model through the communication unit 230 .

[0192] Embodiments of the present disclosure may be implemented as a recording medium including computer executable instructions such as computer executable program modules. Computer readable media may be any available media that can be accessed by a computer, and may include volatile or non-volatile media and removable and non-removable media. In addition, computer readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, or other data. Communication media may typically include computer readable instructions, data structures, or other data of a modulated data signal such as a program module.

[0193] In addition, the computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" refers to a tangible device and does not include a signal (e.g., an electromagnetic wave), and the term "non-transitory storage medium" does not distinguish between a case where data is semi-permanently stored in the storage medium and a case where data is temporarily stored. For example, a non-transitory storage medium may include a buffer in which data is temporarily stored.

[0194] According to an embodiment of the present disclosure, the method according to the embodiment disclosed herein may be included in a computer program product and then provided. The computer program product may be traded between a seller and a buyer as a commodity. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc (CD) ROM (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or distributed directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored in a machine-readable storage medium, such as a memory of a manufacturer's server, an application store's server, or a relay server.

[0195] According to an embodiment of the present disclosure, a method for training a neural network model is provided. The method for training a neural network model includes obtaining a first object recognition result predicted by a first neural network model using a first image captured by a first camera that captures a space including at least one object from a first perspective as an input (S1310). In addition, the method for training a neural network model also includes: converting the predicted first object recognition result based on a conversion relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to a second camera that captures the space from a second perspective (S1320). In addition, the method for training a neural network model also includes: based on the first object recognition result converted relative to the second perspective, generating training data by performing labeling on a second image corresponding to the first image and captured by the second camera (S1330). In addition, the method for training a neural network model also includes training a second neural network model by using the generated training data (S1340).

[0196] In addition, according to an embodiment of the present disclosure, the conversion of the predicted first object recognition result (S1320) includes: based on the conversion relationship, converting the three-dimensional posture information including the spatial coordinates of the feature points representing the position and direction of the object recognized in the first image into the two-dimensional posture information including the image coordinates of the feature points representing the position of the object recognized in the second image. In addition, the generation of training data (S1330) includes generating training data by performing labeling on the second image, so that the two-dimensional posture information included in the first object recognition result converted relative to the second perspective is used as the two-dimensional posture information of the object in the second image.

[0197] In addition, the conversion of the predicted first object recognition result (S1320) further includes: based on the conversion relationship between the first camera coordinate system and the world coordinate system, converting the three-dimensional pose information in the first camera coordinate system into the three-dimensional pose information in the world coordinate system. In addition, the conversion of the predicted first object recognition result (S1320) further includes: based on the conversion relationship between the world coordinate system and the second camera coordinate system, converting the three-dimensional pose information in the world coordinate system into the three-dimensional pose information in the second camera coordinate system. In addition, the conversion of the predicted first object recognition result (S1320) further includes: by projecting the spatial coordinates in the second camera coordinate system to the second image coordinate system of the second camera, converting the three-dimensional pose information in the second camera coordinate system into the two-dimensional pose information in the second image coordinate system.

[0198] In addition, the conversion relationship between the first camera coordinate system and the world coordinate system can be obtained by performing the first camera calibration to obtain the internal parameters of the first camera and the external parameters of the first camera using the matching pairs of the spatial coordinates of a certain number of feature points in the world coordinate system and the image coordinates of the feature points in the image coordinate system of the first camera. In addition, the feature points can be detected by moving the electronic device 100 within the field of view of the first camera and detecting the center point of a certain detection area, a mark attached to the electronic device 100, or a feature of the appearance of the electronic device 100.

[0199] In addition, generating training data (S1330) further includes generating training data by performing labeling on the second image so that, in the two-dimensional posture information included in the first object recognition result converted relative to the second viewing angle, a first portion corresponding to the second image is distinguished from a second portion not corresponding to the second image.

[0200] In addition, according to an embodiment of the present disclosure, when the second camera captures multiple second images from a second perspective, generating training data (S1330) includes: generating training data by performing labeling on each of the multiple second images based on the first object recognition result converted relative to the second perspective.

[0201] In addition, the plurality of second images photographed from the second viewing angle are captured by capturing the dynamic object at preset time intervals in a state where the second camera maintains the second viewing angle.

[0202] In addition, according to an embodiment of the present disclosure, in the case where the second camera captures a plurality of second images from a plurality of different second viewing angles, respectively, the conversion (S1320) includes: for each of the plurality of second viewing angles, based on the conversion relationship between the first camera coordinate system and the second camera coordinate system, converting the predicted first object recognition result. In addition, the generation of training data (S1330) includes: based on the first object recognition result converted with respect to each of the plurality of second viewing angles, generating training data by performing labeling on each of the plurality of second images.

[0203] In addition, the plurality of second images respectively captured from the plurality of different second viewing angles are captured from the plurality of different second viewing angles when the second camera mounted on the electronic device 100 orbits around the static object or changes the distance to the static object.

[0204] In addition, according to an embodiment of the present disclosure, generating training data (S1330) includes generating training data in a case where the electronic device 100 equipped with the second camera recognizes an object at a level lower than a preset standard or moves to a position designated in a space map used by the electronic device 100.

[0205] According to an embodiment of the present disclosure, the electronic device 100 includes a memory 110, a processor 120 configured to execute at least one instruction stored in the memory 110, a camera 131, and a communication unit 140. In addition, the processor 120 executes at least one instruction to obtain, through the communication unit 140, a first object recognition result predicted by a first neural network model using a first image captured from a first perspective by a camera of an external device 300 that captures a space including at least one object as an input. In addition, the processor 120 executes at least one instruction to convert the predicted first object recognition result based on a conversion relationship between a first camera coordinate system corresponding to the camera of the external device 300 and a second camera coordinate system corresponding to the camera 131 that captures the space from a second perspective. In addition, the processor 120 executes at least one instruction to generate training data by performing labeling on a second image corresponding to the first image and captured by the camera 131 based on the first object recognition result converted relative to the second perspective. In addition, the processor 120 executes at least one instruction to train the second neural network model by using the generated training data.

[0206] In addition, according to an embodiment of the present disclosure, the processor 120 executes at least one instruction to convert the three-dimensional pose information into two-dimensional pose information based on the conversion relationship, the three-dimensional pose information includes the spatial coordinates of the feature points representing the position and direction of the object recognized in the first image, and the two-dimensional pose information includes the image coordinates of the feature points representing the position of the object recognized in the second image. In addition, the processor 120 executes at least one instruction to generate training data by performing labeling on the second image, so that the two-dimensional pose information included in the first object recognition result converted relative to the second perspective is used as the two-dimensional pose information of the object in the second image.

[0207] In addition, the processor 120 executes at least one instruction to convert the three-dimensional pose information in the first camera coordinate system into the three-dimensional pose information in the world coordinate system based on the conversion relationship between the first camera coordinate system and the world coordinate system. In addition, the processor 120 executes the at least one instruction to convert the three-dimensional pose information in the world coordinate system into the three-dimensional pose information in the second camera coordinate system based on the conversion relationship between the world coordinate system and the second camera coordinate system. In addition, the processor 120 executes at least one instruction to convert the three-dimensional pose information in the second camera coordinate system into the two-dimensional pose information in the second image coordinate system by projecting the spatial coordinates in the second camera coordinate system into the second image coordinate system of the camera.

[0208] In addition, the processor 120 executes the at least one instruction to obtain a conversion relationship between the first camera coordinate system and the world coordinate system by performing the first camera calibration, thereby obtaining the internal parameters of the camera of the external device 300 and the external parameters of the camera of the external device 300 using matching pairs of spatial coordinates of a certain number of feature points in the world coordinate system and image coordinates of feature points in the image coordinate system of the camera of the external device 300. In addition, the feature points are detected by moving the electronic device 100 within the field of view of the camera of the external device 300 and detecting a center point of a certain detection area, a marker attached to the electronic device 100, or a feature of the appearance of the electronic device 100.

[0209] In addition, the processor 120 executes the at least one instruction to generate training data by performing labeling on the second image, so that in the two-dimensional posture information included in the first object recognition result converted relative to the second perspective, the first part corresponding to the second image is distinguished from the second part that does not correspond to the second image.

[0210] In addition, according to an embodiment of the present disclosure, the processor 120 executes at least one instruction to generate training data by labeling each of the multiple second images based on the first object recognition result converted relative to the second perspective when the camera 131 captures multiple second images from a second perspective.

[0211] In addition, according to an embodiment of the present disclosure, the processor 120 executes at least one instruction to transform the predicted first object recognition result based on the transformation relationship between the first camera coordinate system and the second camera coordinate system for each of the multiple second perspectives when the camera 131 captures multiple second images from multiple different second perspectives, and generates training data by performing labeling on each of the multiple second images based on the first object recognition result transformed relative to each of the multiple second perspectives.

[0212] In addition, according to an embodiment of the present disclosure, in a case where the electronic device 100 recognizes an object at a level below a preset standard or moves to a position specified in a space map used by the electronic device 100 , the processor 120 executes at least one instruction to generate training data.

[0213] According to an embodiment of the present disclosure, the cloud server 200 includes a memory 210, a processor 220 configured to execute at least one instruction stored in the memory 210, and a communication unit 230. The processor 220 executes at least one instruction to obtain a first object recognition result predicted by a first neural network model using a first image captured by a first camera capturing a space including at least one object from a first perspective as an input. In addition, the processor 220 executes at least one instruction to obtain a second image captured by a second camera capturing the space from a second perspective through the communication unit 230. In addition, the processor 220 executes the at least one instruction to transform the predicted first object recognition result based on a transformation relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to the second camera. In addition, the processor 220 executes at least one instruction to generate training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective. In addition, the processor 220 executes at least one instruction to train the second neural network model by using the generated training data.

[0214] According to an embodiment of the present disclosure, the processor 220 executes at least one instruction to send the trained network parameter value of the second neural network model to the electronic device 100 equipped with the second neural network model through the communication unit 230.

[0215] Although the present disclosure has been specifically shown and described, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present disclosure. Therefore, it should be understood that the above embodiments are exemplary in all respects and do not limit the scope of the present disclosure. For example, each element described in a single type may be performed in a distributed manner, and the elements described in a distributed manner may also be performed in an integrated form.

[0216] The scope of the present disclosure is not limited by the detailed description of the present disclosure but by the appended claims, and all modifications or substitutions derived from the scope and spirit of the claims and their equivalents fall within the scope of the present disclosure.

Claims

1. A computer-implemented method for training a neural network model by augmenting images representing an object captured by a plurality of cameras, the method include: obtaining a first object recognition result predicted by a first neural network model using a first image captured by a first camera as an input, the first camera capturing a space including at least one object from a first perspective (S1310); converting the obtained first object recognition result based on a conversion relationship between a first camera coordinate system corresponding to the first camera and a second camera coordinate system corresponding to a second camera capturing the space from a second perspective (S1320); generating training data by performing labeling on a second image corresponding to the first image based on the first object recognition result converted with respect to a second viewing angle, the second image being captured by the second camera (S1330); and The second neural network model is trained by using the generated training data ( S1340 ).

2. The computer-implemented method of claim 1, in, The conversion of the obtained first object recognition result (S1320) includes: based on the conversion relationship, converting the three-dimensional posture information into two-dimensional posture information, The three-dimensional posture information includes spatial coordinates of feature points representing the position and orientation of the object identified in the first image, wherein the two-dimensional posture information includes image coordinates of feature points representing positions of objects identified in the second image, The generating of training data (S1330) includes generating training data by performing labeling on the second image so that two-dimensional pose information included in the first object recognition result converted with respect to the second viewing angle is used as two-dimensional pose information of the object in the second image.

3. A computer-implemented method according to claim 1 or claim 2, in, Converting the obtained first object recognition result (S1320) further includes: Based on a conversion relationship between the first camera coordinate system and a world coordinate system, converting the three-dimensional posture information in the first camera coordinate system into the three-dimensional posture information in the world coordinate system; Based on a conversion relationship between the world coordinate system and the second camera coordinate system, converting the three-dimensional posture information in the world coordinate system into the three-dimensional posture information in the second camera coordinate system; and The three-dimensional pose information in the second camera coordinate system is converted into two-dimensional pose information in the second image coordinate system by projecting the spatial coordinates in the second camera coordinate system to the second image coordinate system of the second camera.

4. A computer-implemented method according to any one of claims 1 to 3, in, The conversion relationship between the first camera coordinate system and the world coordinate system is obtained by performing a first camera calibration to obtain an intrinsic parameter of the first camera and an extrinsic parameter of the first camera using matching pairs of spatial coordinates of a plurality of feature points in the world coordinate system and image coordinates of feature points in the image coordinate system of the first camera, and The feature point is detected by moving the electronic device (100) within the field of view of the first camera and detecting a center point of a detection area, a mark attached to the electronic device (100), or an appearance of the electronic device (100).

5. A computer-implemented method according to any one of claims 1 to 4, in, Generating the training data (S1330) also includes: generating the training data by performing labeling on the second image, so that, among the two-dimensional posture information included in the first object recognition result converted relative to the second viewing angle, a first part corresponding to the second image is distinguished from a second part not corresponding to the second image.

6. A computer-implemented method according to any one of claims 1 to 5, in, Generating the training data based on a plurality of second images captured by the second camera from the second viewing angle (S1330) includes generating the training data by performing labeling on at least one second image of the plurality of second images based on the first object recognition result converted relative to the second viewing angle.

7. A computer-implemented method according to any one of claims 1 to 6, in, The plurality of second images captured from the second viewing angle are captured by capturing the dynamic object at preset time intervals, and Wherein the second camera is configured to maintain the second viewing angle.

8. A computer-implemented method according to any one of claims 1 to 7, in, Based on a plurality of second images captured by the second camera from a plurality of different second viewing angles respectively, the converting (S1320) includes: for at least one of the plurality of second viewing angles, based on a conversion relationship between the first camera coordinate system and the second camera coordinate system, converting the obtained first object recognition result, The generating of the training data (S1330) includes: generating the training data by labeling at least one of the plurality of second images based on the first object recognition result converted relative to at least one of the plurality of second perspectives.

9. A computer-implemented method according to any one of claims 1 to 8, in, The plurality of second images captured respectively from the plurality of different second viewing angles are captured from the plurality of different second viewing angles, and The second camera installed on the electronic device (100) is configured to orbit around a static object or change the distance to the static object.

10. A computer-implemented method according to any one of claims 1 to 9, in, Generating the training data (S1330) includes generating the training data based on an electronic device (100) equipped with the second camera, the second camera being configured to recognize an object at a level lower than a preset standard or being configured to move to a position specified in a spatial map used by the electronic device (100).

11. An electronic device (100), include: Camera (131); Communication unit (140); A memory (110) configured to store at least one instruction; At least one processor (120) is operably connected to the camera (131), the communication unit (140) and the memory (110), and is configured to execute the at least one instruction to: obtaining, through the communication unit (140), a first object recognition result predicted by a first neural network model, the first neural network model using as input a first image captured from a first perspective by a camera of an external device (300) capturing a space including at least one object, converting the obtained first object recognition result based on a conversion relationship between a first camera coordinate system corresponding to a camera of the external device (300) and a second camera coordinate system corresponding to a camera (131) capturing the space from a second perspective, generating training data by performing labeling on a second image corresponding to the first image based on the first object recognition result transformed relative to the second perspective, the second image being captured by the camera (131); and The second neural network model is trained by using the generated training data.

12. The electronic device (100) according to claim 11, in, The at least one processor (120) is further configured to execute the at least one instruction to: Based on the conversion relationship, the three-dimensional posture information is converted into two-dimensional posture information, The three-dimensional posture information includes spatial coordinates of feature points representing the position and orientation of the object identified in the first image, and wherein the two-dimensional posture information includes image coordinates of feature points representing positions of objects identified in the second image, and The training data is generated by performing labeling on the second image so that two-dimensional pose information included in the first object recognition result converted with respect to the second viewing angle is used as two-dimensional pose information of the object in the second image.

13. The electronic device (100) according to claim 11 or 12, in, The at least one processor (120) is also configured to execute the at least one instruction to: generate the training data based on a plurality of second images captured by the camera (131) from the second perspective, by performing labeling on at least one second image of the plurality of second images based on the first object recognition result converted relative to the second perspective.

14. The electronic device (100) according to any one of claims 11 to 13, in, The at least one processor (120) is further configured to execute the at least one instruction to: Based on a plurality of second images captured by the camera (131) from a plurality of different second viewing angles respectively, and based on a conversion relationship between the first camera coordinate system and the second camera coordinate system for each of the plurality of second viewing angles, converting the obtained first object recognition result; as well as The training data is generated by performing labeling on at least one of the plurality of second images based on the first object recognition result transformed with respect to each of the plurality of second perspectives.

15. The electronic device (100) according to any one of claims 11 to 14, in, The at least one processor (120) is further configured to execute the at least one instruction to generate the training data, and The electronic device (100) is configured to identify an object at a level lower than a preset standard, or to move to a position specified in a spatial map used by the electronic device (100).