Method and device for generating artificial intelligence model for recognizing personalized object, and method and system for controlling robot using same
An AI model for robots recognizes personalized objects using user environment images, addressing the challenge of intuitive human-robot interactions by enabling user-centric command recognition.
Patent Information
- Application Number
- PCT/KR2024/012751
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2024-08-27
- Publication Date
- 2026-01-02
Smart Images

Figure KR2024012751_02012026_PF_FP_ABST
Abstract
Description
Method and device for creating an artificial intelligence model for recognizing personalized objects, and a robot control method and system using the same
[0001] The present invention relates to a method and device for training an object-grabbing robot, and more particularly, to a method and system for generating an artificial intelligence model that recognizes a personalized object by utilizing images collected from a user's environment, and controlling a robot configured to recognize and grasp an object by utilizing the artificial intelligence model.
[0002] Empowering robots to understand human natural language is a crucial challenge in the fields of AI and robotics. If robots can follow and execute human natural language instructions, more intuitive human-robot interactions are possible.
[0003] To enable robots to understand human natural language, research has traditionally focused on the Language-Conditioned Robotic Grasping (LCRG) approach. The LCRG approach focuses on robotic systems that grasp objects based on verbal instructions.
[0004] The LCRG approach relies on general language expressions rather than personalized expressions when describing objects, which inevitably leads to less intuitive human-robot interactions. For example, users must be more specific, such as "bring me the black wallet next to my laptop" rather than "bring me my wallet."
[0005] This type of robot-centric, rather than user-centric, instruction presents a problem: it's cumbersome for many users. Furthermore, for more personalized interactions with robots, more personalized data must be trained on the robots.
[0006] Accordingly, there is a need for a method for creating an artificial intelligence model that enables a robot to identify objects with just one interaction with a human and recognize personalized objects based on natural language instructions, and a method for controlling the robot using the model.
[0007] In order to solve the above-described technical problem, the present invention provides a method and device for generating an artificial intelligence model that recognizes a personalized object of a specific user included in an image collected in a preset user environment.
[0008] In addition, in order to solve the above-described technical problem, the present invention seeks to provide a robot control method and system using an artificial intelligence model that recognizes a personalized object.
[0009] The technical problems to be solved by the present invention are not limited to the technical problems described above, and other technical problems of the present invention can be derived from the following description.
[0010] In order to solve the above-described technical problem, one embodiment of the present invention provides a method for generating an artificial intelligence model for recognizing a personalized object, which is performed by at least one processor. The method includes the steps of: receiving object information about a personalized object of a specific user included in a first image collected in a preset user environment and labeling the personalized first object of the specific user; assigning a label identical to the label of the first object to a second object having a similarity greater than or equal to a preset value of the first object based on a similarity between an embedding vector of the first object for which labeling has been completed and an embedding vector of a second object included in a second image collected in the preset user environment; and using the first image and the second image as a learning data set and including a labeling process for the first object and the second object, and generating an artificial intelligence model trained to output a control signal according to a command when a command for a personalized object of a specific user is received.
[0011] In addition, another embodiment of the present invention provides an artificial intelligence model generating device for recognizing a personalized object. The device includes a communication module, at least one processor, and a memory electrically connected to the processor and storing at least one code to be executed by the processor. The memory stores code that, when executed by the processor, causes the processor to receive object information regarding a personalized object of a specific user included in a first image collected from a user environment, perform labeling on the personalized first object of the specific user, and assign a label identical to the label of the first object to a second object having a similarity greater than or equal to a preset value based on a similarity between an embedding vector of the first object for which labeling has been completed and an embedding vector of a second object included in a second image collected from a preset user environment, and uses the first image and the second image as a learning data set, and includes a labeling process for the first object and the second object, and generates an artificial intelligence model that is trained to output a control signal according to the command when a command regarding the personalized object of the specific user is received.
[0012] Another embodiment of the present invention provides a robot control method using an artificial intelligence model. The method comprises the steps of receiving a command from a user regarding an object of the user, searching for an object identical to the user's object among objects collected from images in a preset user environment, and controlling the robot based on a control signal output by an artificial intelligence model that recognizes the personalized object.
[0013] In addition, another embodiment of the present invention provides a robot control system using an artificial intelligence model. The system includes a communication module, at least one processor, and a memory electrically connected to the processor and storing at least one code executed by the processor. The memory stores code that, when executed by the processor, causes the processor to receive a command from a user regarding an object of the user, search for an object identical to the user's object among objects collected from images in a preset user environment, and control the robot based on a control signal output by an artificial intelligence model that recognizes the personalized object.
[0014] According to the solution to the problem of the present invention described above, robot learning can be performed without requiring additional effort from the user.
[0015] In addition, according to the solution to the problem of the present invention described above, the robot can be controlled through user-centered natural language instructions rather than robot-centered instructions.
[0016] In addition, according to the solution to the problem of the present invention described above, it is possible to identify a user's personal object without relying on the user's supervised learning.
[0017] In addition, according to the solution to the problem of the present invention described above, the appearance of an object viewed from various angles can be effectively utilized for learning through robot-object interaction that captures the object from various angles.
[0018] The effects of the present invention are not limited to the effects described above, and include all effects understood from the following description.
[0019] FIG. 1 is a drawing illustrating a server and a robot connected to the server in communication with the server according to one embodiment of the present invention.
[0020] Figure 2 is a drawing showing the detailed configuration of the server illustrated in Figure 1.
[0021] Figure 3 is a diagram illustrating in detail the process by which the server illustrated in Figure 1 trains an artificial intelligence model.
[0022] Figure 4 is a diagram showing the results of an object grasping experiment using the robot illustrated in Figure 1, classified by type of input data.
[0023] FIG. 5 is a flowchart illustrating the sequence of a method for generating an artificial intelligence model that recognizes a personalized object according to another embodiment of the present invention.
[0024] Figure 6 is a flowchart illustrating the sequence of a robot control method using an artificial intelligence model according to another embodiment of the present invention.
[0025] Hereinafter, the present invention will be described in detail with reference to the attached drawings. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. In addition, the attached drawings are only intended to facilitate understanding of the embodiments of the invention disclosed herein, and the technical ideas disclosed herein are not limited by the attached drawings. All terms, including technical and scientific terms, used herein should be interpreted as having meanings generally understood by a person of ordinary skill in the art to which the present invention pertains. Terms defined in the dictionary should be interpreted as having additional meanings consistent with the relevant technical literature and the present invention, and shall not be interpreted in an extremely ideal or restrictive sense unless otherwise defined.
[0026] In order to clearly explain the present invention in the drawings, parts irrelevant to the description have been omitted, and the size, shape, and appearance of each component shown in the drawings may be modified in various ways. Identical / similar parts throughout the specification are given identical / similar drawing reference numerals.
[0027] The suffixes "module" and "function" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. Furthermore, in describing the embodiments disclosed herein, detailed descriptions of related known technologies have been omitted if they are deemed to obscure the gist of the embodiments disclosed herein.
[0028] Throughout the specification, when a part is said to be "connected (connected, in contact with, or coupled)" to another part, this includes not only cases where it is "directly connected (connected, in contact with, or coupled)" but also cases where it is "indirectly connected (connected, in contact with, or coupled)" with another member in between. Furthermore, when a part is said to "include (have or provide)" a certain component, this does not mean that it excludes other components, but rather that it may "include (have or provide)" other components, unless otherwise specifically stated.
[0029] As used herein, ordinal terms such as "first," "second," etc., are used solely to distinguish one component from another and do not limit the order or relationship of the components. For example, the first component of the present invention may be referred to as the "second component," and similarly, the second component may also be referred to as the "first component." As used herein, singular expressions should be construed to include plural expressions, unless explicitly stated otherwise.
[0030] FIG. 1 is a drawing illustrating a server and a robot connected to the server in communication with the server according to one embodiment of the present invention.
[0031] In one example, the server (100) may be at least one of an artificial intelligence model generation device that recognizes a personalized object (hereinafter referred to as an “artificial intelligence model generation device”) and a robot control device that uses an artificial intelligence model that recognizes a personalized object (hereinafter referred to as a “robot control device”).
[0032] The artificial intelligence model generation device can be connected to a robot (200). The artificial intelligence model generation device can receive images from the robot (300). However, the image reception is not limited to the robot (300), and the image can be received from at least one of a storage device storing at least one image or a photographing device including a camera connected to a network.
[0033] The robot control device can be connected to the robot (200) for communication. The robot control device can receive images from the robot (300). However, the image reception is not limited to the robot (300), and the image can be received from at least one of a storage device storing at least one image or a photographing device including a camera connected to a network.
[0034] Figure 2 is a drawing showing the detailed configuration of the server illustrated in Figure 1.
[0035] Referring to FIG. 2, the server (100) may be at least one of an artificial intelligence model generation device and a robot control device. In this case, the artificial intelligence model generation device and the robot control device may be referred to as an artificial intelligence model generation system and a robot control system, respectively.
[0036] The artificial intelligence model generation device may include a communication module (110), a processor (120) that performs operations according to code stored in a memory (130), and a memory (130) that stores the code.
[0037] The AI model generation device can be implemented as a computer or portable terminal that can connect to a server or other terminal via a network. Here, the computer includes, for example, a notebook, desktop, or laptop equipped with a web browser, and the portable terminal can include, for example, a wireless communication device that guarantees portability and mobility, and can include all types of handheld-based wireless communication devices such as various types of communication-based terminals, smartphones, and tablet PCs. In addition, the portable terminal can be an edge device or on-device AI having at least one processor capable of AI model inference tasks. The network can be implemented as a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or any type of wireless network such as a mobile radio communication network or satellite communication network.
[0038] In the artificial intelligence model generation device, the communication module (110) may include a device including hardware and software necessary for transmitting and receiving signals such as control signals or data signals through wired or wireless connections with other network devices. In a mobile communication network constructed according to technical standards or communication methods for mobile communication used in the mobile communication module (e.g., GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTEA (Long Term Evolution-Advanced), etc.), a wireless signal is transmitted and received with at least one of a base station, an external terminal, and a server.
[0039] In the artificial intelligence model generation device, the processor (120) may include various types of devices that control and process data. The processor (120) may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program. In one example, the processor (120) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.
[0040] In the artificial intelligence model generation device, the memory (130) can store at least one of information and data input to the communication module (110), information and data required for a function performed by the processor (120), and data generated according to the execution of the processor (120).
[0041] In the artificial intelligence model generation device, the memory (130) should be interpreted as a general term for a non-volatile storage device that maintains stored information even when no power is supplied and a volatile storage device that requires power to maintain the stored information. In addition to a volatile storage device that requires power to maintain the stored information, the memory (130) may include cloud storage, SSD, magnetic storage media, or flash storage media, but the scope of the present invention is not limited thereto.
[0042] In the artificial intelligence model generation device, the memory (130) is electrically connected to the processor (120) and stores at least one code that is executed by the processor (120). The memory (130) stores a code that causes the processor (120) to perform the following functions and procedures when executed by the processor (120).
[0043] In an artificial intelligence model generation device, a memory (130) stores code that causes a processor (120) to receive object information about a personalized object of a specific user included in a first image collected in a preset user environment, perform labeling on the personalized first object of the specific user, assign a label identical to the label of the first object to a second object having a preset similarity or higher to the first object based on a similarity between an embedding vector of the first object for which labeling has been completed and an embedding vector of a second object included in a second image collected in a preset user environment, and use the first image and the second image as a learning data set, and include a labeling process for the first object and the second object, and generate an artificial intelligence model trained to output a control signal according to the command when a command for the personalized object of the specific user is received. The object information may include a natural language instruction including ownership information for the first object and images of the first object from various angles.
[0044] In the artificial intelligence model generation device, the memory (130) may further store code that causes the processor (120) to repeatedly assign the same label as the label of the first object to a second object having a similarity greater than or equal to a preset value of the first object until the proportion of objects to which labels are not assigned is less than 10%. The similarity may be calculated based on the cosine similarity between the embedding vector of the first object, for which labeling has been completed, and the embedding vector of the second object.
[0045] The robot control device may include a communication module (110), a processor (120) that performs operations according to code stored in a memory (130), and a memory (130) that stores the code.
[0046] The robot control device can be implemented as a computer or portable terminal that can connect to a server or other terminal via a network. Here, the computer includes, for example, a notebook, desktop, or laptop equipped with a web browser, and the portable terminal can include, for example, a wireless communication device that guarantees portability and mobility, and can include all types of handheld-based wireless communication devices such as various types of communication-based terminals, smartphones, and tablet PCs. In addition, the portable terminal can be an edge device or on-device AI having at least one processor capable of AI model inference. The network can be implemented as a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or any type of wireless network such as a mobile radio communication network or satellite communication network.
[0047] In the robot control device, the communication module (110) may include a device including hardware and software necessary for transmitting and receiving signals such as control signals or data signals through a wired or wireless connection with other network devices. In the mobile communication module, a wireless signal is transmitted and received with at least one of a base station, an external terminal, and a server on a mobile communication network constructed according to technical standards or communication methods for mobile communication (e.g., GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTEA (Long Term Evolution-Advanced), etc.).
[0048] In a robot control device, the processor (120) may include various types of devices that control and process data. The processor (120) may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program. In one example, the processor (120) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.
[0049] In the robot control device, the memory (130) can store at least one of information and data input to the communication module (110), information and data required for a function performed by the processor (120), and data generated according to the execution of the processor (120).
[0050] In a robot control device, memory (130) should be interpreted as a general term for a non-volatile storage device that maintains stored information even when no power is supplied and a volatile storage device that requires power to maintain the stored information. In addition to a volatile storage device that requires power to maintain the stored information, memory (130) may include cloud storage, SSD, magnetic storage media, or flash storage media, but the scope of the present invention is not limited thereto.
[0051] In the robot control device, the memory (130) is electrically connected to the processor (120) and stores at least one code that is executed by the processor (120). The memory (130) stores a code that causes the processor (120) to perform the following functions and procedures when executed by the processor (120).
[0052] In a robot control device, a memory (130) stores a code that causes a processor (120) to receive a command from a user regarding an object of the user, search for an object identical to the user's object among objects in images collected in a preset user environment, and control the robot based on a control signal output by an artificial intelligence model that recognizes a personalized object. The artificial intelligence model that recognizes a personalized object may be an artificial intelligence model that uses images collected in a preset user environment and images of the user's object as a learning data set, includes a labeling process for objects in the images collected in the preset user environment and the user's object, and is trained to output a control signal according to the command when a command regarding the user's object is received.
[0053] In the robot control device, the memory (130) may further store code that causes the processor (120) to infer bounding box coordinates for the user's object based on a command. The command may include ownership information for the user's object.
[0054] In the robot control device, the memory (130) can further store a code that causes the processor (120) to detect the user's object, obtain three-dimensional coordinates of the user's object through a preset algorithm, and manipulate the object.
[0055] Figure 3 is a diagram illustrating in detail the process by which the server illustrated in Figure 1 trains an artificial intelligence model.
[0056] The server (100) may be at least one of an artificial intelligence model generation device and a robot control device.
[0057] The artificial intelligence model generation device is located in the user environment. Collect images of dogs and call them Reminiscences. (S10). The artificial intelligence model generation device detects all objects in the memory through an off-the-shelf object detector. Detects and creates a node for each detected object, and the node is a pre-trained visual encoder. and DINO (Self-Distilation with no Labels) are used for embedding. Embedded object nodes are a set of vector embeddings. It appears as, If The bounding box (bbox) coordinates of the detected images and objects are tagged to the embedded object nodes. At this time, There are no private indicators of objects in the node.
[0058] The artificial intelligence model generation device obtains object information (S20) based on human-robot interaction (S21) and robot-object interaction (S22).
[0059] In human-robot interaction (S21), the user has his / her own personal object and describe it in natural language. The natural language description includes general indicators. and personal indicators Includes general indicators For example, it may contain state information or location information of an object, such as 'object in front', and personal indicators For example, an object may include ownership information or the name of the object, such as 'my sleeping pills'. The artificial intelligence model generation device may include visual recognition indicators. and general indicators Using the bounding box coordinates of an object Predicting and pre-trained visual encoders Using object nodes Creates. This node contains , and Meta information including is labeled.
[0060] In robot-object interaction (S22), the artificial intelligence model generation device The object-grabbing robot is made to grab the object and then rotate it. At this time, the artificial intelligence model generation device Obtain an image of the object and obtain images of the object from multiple angles. The artificial intelligence model generation device acquires the acquired image A set of objects in and bounding box coordinates of sets of objects extracts. The server (100) is a set of object nodes , and at this time each node In , and Meta information including is labeled.
[0061] After human-robot interaction (S21) and robot-object interaction (S22), each individual object and , that is, personal indicators are tagged It is connected to the nodes of the dog. Among all individual objects, the set of nodes labeled with individual indicators is , and a set of personal indicators is formed. It forms.
[0062] The artificial intelligence model generation device By propagating the label in Label the missing indicators (S30). In more detail, Each node of About, Affinity score for , and the indicator with the highest affinity score is calculated based on the affinity score. to node is assigned. At this time, it is assigned only when the highest affinity score is above a certain threshold. The node with the most likely personal indicator After labeling Is , and label propagation is repeated until the proportion of relabeled nodes is less than 10%.
[0063] Affinity Score Is Nodes and Indicators All object nodes tagged with Average cosine similarity between is. Affinity score is calculated from the following [Mathematical Formula 1], and the average cosine similarity is calculated from the following [Mathematical Formula 2]. is an indicator Tagged with It represents the number of nodes. [Mathematical formula 1] is an indicator function that has 1 if the condition is true and 0 otherwise. In addition, , am.
[0064]
[0065]
[0066] The process of training an artificial intelligence model by the server (100) illustrated in Fig. 1 follows the following algorithm.
[0067] The set of user's personalized objects
[0068] Visual Encorder
[0069] Nodes from
[0070] Set of personal indicators
[0071] Nodes with indicators
[0072] for do
[0073] and from human-robot interaction
[0074] Get from robot-object interaction
[0075] tagged with
[0076] and
[0077] end for
[0078]
[0079] for do
[0080] , set
[0081] for do
[0082]
[0083] and
[0084] end for
[0085]
[0086] if max then
[0087]
[0088]
[0089] end if
[0090] end for
[0091] until Ratio of changed indicator of node < 10%
[0092] An AI model that recognizes personalized objects can be a Transformer-based vision-language model. The personalized object recognition model can infer the bounding box coordinates of a personalized object through natural language instructions after receiving an image. Optimization of the personalized object recognition model is based on the following mathematical equation (3) for each object node. This is done by minimizing the negative log probability for the bounding box of . At this time, is a node after the step of attaching labels to missing indicators (S30).
[0093]
[0094] The process of grasping an object using a robot (300) begins when a user instructs the robot (300) to grasp a personalized object. The robot control device receives the user's command and estimates the 2D bounding box coordinates of the instructed object using an artificial intelligence model that recognizes the personalized object, and then calculates the 3D coordinates of the object using point cloud data and the RANSAC algorithm. More specifically, the point cloud is used to convert the 2D bounding box coordinates into a 3D spatial structure, and then the RANSAC algorithm is used to segment the object points within the 3D bounding box coordinates. Thereafter, the robot control device calculates the path of the robot's (300) arm so that the robot (300) grasps the object.
[0095] Figure 4 is a diagram illustrating the results of an object grasping experiment using the robot depicted in Figure 1, categorized by input data type. The present invention is further described below using examples. The examples were evaluated using the following tests.
[0096] The "PGA" described below is an artificial intelligence model substantially identical to the artificial intelligence model for recognizing personalized objects of the present invention. In other words, the operations and functions of the PGA described below can be performed by the artificial intelligence model for recognizing personalized objects.
[0097] The AI model for recognizing personalized objects of the present invention used in each embodiment comprises a curated dataset comprising a training set, a recall set, and a test set. The curated dataset includes over 100 general objects and 96 personalized objects.
[0098] The training set contains each individual indicator 96 pairs of images obtained from human-robot interactions Reminiscence contains 400 images containing multiple objects. Unlabeled images, such as the above images, can be optionally provided during the training process. The training set includes Heterogeneous split, Homogeneous split, Cluttered split, Paraphrased split, and Generic split.
[0099] The heterogeneous split contains scenes containing randomly selected objects, with 60 images and 120 individual landmarks and bounding box coordinates. The homogeneous split contains scenes containing similar objects from the same category, with 60 images and 120 individual landmarks and bounding box coordinates. The cluttered split contains images containing cluttered objects, each with a single individual landmark and bounding box coordinates. Each image is taken from the IM-Dial dataset. The paraphrased split contains all heterogeneous, homogeneous, and cluttered splits containing individual landmarks. The generic split is taken from the dataset provided by the GVCCI.
[0100] In this embodiment, the target object grounding accuracy is evaluated using the Interaction over Union (IoU) score when individual metrics are given. IoU is a factor that calculates the overlap between the predicted and actual bounding box coordinates. In Example 1, while presenting a prediction rate with an IoU exceeding 0.5, an IoU score exceeding 0.8 is recorded to ensure the precision required for successful object grasping. The evaluation results of Example 1 are presented in [Table 1] below.
[0101] HeterogeneousHomogeneousParaphrasedClutteredGenericMethodinteractionannotatedutilizedIoU>0.5IoU>0.8IoU>0.5IoU>0.8IoU>0.5IoU>0.8IoU>0.8IoU>0.5OFA---49.247.523.720.335.834.134.665.7GVCCI---59.355.130.523.842.338.944.379.1Direct1.1969660.255.137.327.148.342.546.4-PassivePGA1.196482889.076.368.653.473.662.762.4-PGA(ours)1.196649291.581.470.361.974.768.264.579.1Supervised91.38763876397.590.792.483.184.879.071.4-
[0102] The experimental results presented in Table 1 show that the Direct model, which relies on object information obtained from human-robot interactions without reminiscence, only slightly improves performance compared to methods that rely on general object knowledge (e.g., OFA and GVCCI). This demonstrates that training the LCRG model with only a small amount of labeled data does not significantly improve performance. The AI model for recognizing personalized objects, utilizing a large amount of unlabeled data, showed approximately 30% better performance than the Direct model. Notably, compared to the Supervised model, the Supervised model showed even better results, despite having approximately 100 times more annotations than the AI model for recognizing personalized objects. We evaluated the performance of the PGA when querying with generic metrics using generic split to determine how well the PGA retains its knowledge of generic metrics after training with personalized metrics. The acquisition of personalized metrics did not impair the PGA's ability to recognize generic metrics. In summary, by leveraging reminiscence and the robot's manipulation capabilities, we were able to efficiently ground learned individual objects in a single human-robot interaction, maintaining knowledge of general representations while maintaining performance comparable to supervised approaches.
[0103] Referring to Figure 4, this embodiment can confirm the relationship between PGA performance and the size of the recall, i.e., the number of raw images used. Experiments were conducted at sizes of 25, 100, and 400 in a logarithmic scale, and the results show that PGA performance is proportional to the size of the recall, i.e., the number of raw images used.
[0104] FIG. 5 is a flowchart illustrating the sequence of a method for generating an artificial intelligence model for recognizing a personalized object according to another embodiment of the present invention.
[0105] The object-grabbing robot learning method described below can be performed by the artificial intelligence model generation device described above with reference to FIGS. 1 to 4. Therefore, the content of the embodiments of the present invention described above with reference to FIGS. 1 to 4 can be equally applied to the embodiments described below, and any content overlapping with the above description will be omitted. The steps described below do not necessarily have to be performed in order, and the order of the steps can be set in various ways, and the steps can be performed almost simultaneously.
[0106] Referring to FIG. 5, a method for generating an artificial intelligence model for recognizing a personalized object includes a step of receiving first object information and labeling the first object (S110), a step of assigning the same label as the first object to a second object (S120), and a step of generating an artificial intelligence model (S130).
[0107] The step of receiving first object information and labeling the first object (S110) is a step of receiving object information about a personalized object of a specific user included in a first image collected in a preset user environment, and performing labeling on the personalized first object of the specific user. The object information may include natural language instructions containing ownership information about the first object and multi-angle images of the first object.
[0108] The step (S120) of assigning the same label to the second object as to the first object is a step of assigning the same label to the second object that has a similarity greater than or equal to a preset value between the embedding vector of the first object, for which labeling has been completed, and the embedding vector of the second object included in the second image collected in a preset user environment. The step (S120) of assigning the same label to the second object as to the first object may be repeated until the proportion of objects to which labels have not been assigned is less than 10%. The similarity may be calculated based on the cosine similarity between the embedding vector of the first object, for which labeling has been completed, and the embedding vector of the second object.
[0109] The artificial intelligence model creation step (S130) is a step of creating an artificial intelligence model that is trained to output a control signal according to a command when a command for a personalized object of a specific user is received, including a labeling process for the first object and the second object, using the first image and the second image as a learning data set.
[0110] Figure 6 is a flowchart illustrating the sequence of a robot control method using an artificial intelligence model according to another embodiment of the present invention.
[0111] The robot control method using the artificial intelligence model described below can be performed by the robot control device described above with reference to FIGS. 1 to 4. Accordingly, the content of the embodiments of the present invention described above with reference to FIGS. 1 to 4 can be equally applied to the embodiments described below, and any content overlapping with the above description will be omitted. The steps described below do not necessarily have to be performed in order, and the order of the steps can be set in various ways, and the steps can be performed almost simultaneously.
[0112] Referring to FIG. 6, a robot control method using an artificial intelligence model includes a command receiving step for an object (S210), an identical object search step (S220), and a robot control step (S230).
[0113] The step of receiving a command for an object (S210) is a step of receiving a command for the user's object from the user. The step of receiving a command for an object (S210) may include a step of inferring bounding box coordinates for the user's object based on the command for the object. The command may include ownership information for the user's object.
[0114] The same object search step (S220) is a step of searching for an object that is identical to the user's object among objects in images collected in a preset user environment.
[0115] The robot control step (S230) is a step of controlling the robot based on a control signal output by an artificial intelligence model that recognizes a personalized object. The robot control step (S230) may be a step in which the robot detects the user's object, obtains three-dimensional coordinates of the user's object through a preset algorithm, and manipulates the object. The artificial intelligence model that recognizes the personalized object may be an artificial intelligence model that uses images collected from a preset user environment and images of the user's object as a learning data set, includes a labeling process for objects in the images collected from the preset user environment and the user's object, and is trained to output a control signal according to the command when a command for the user's object is received.
[0116] The method for generating an artificial intelligence model for recognizing a personalized object of the present invention described so far can also be implemented in the form of a recording medium containing computer-executable instructions, such as a program module executed by a computer. The computer-readable medium may be any available medium that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. In addition, the computer-readable medium may include a computer storage medium. The computer storage medium includes both volatile and nonvolatile, removable and non-removable media implemented with any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.
[0117] Those skilled in the art will appreciate that the present invention can be easily modified into other specific forms without changing the technical spirit or essential characteristics of the present invention based on the above description. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. The scope of the present invention is indicated by the following claims, and all changes or modifications derived from the meaning and scope of the claims and their equivalents should be interpreted as being included in the scope of the present invention. The scope of the present application is indicated by the following claims rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalents should be interpreted as being included in the scope of the present application.
Claims
1. A method for generating an artificial intelligence model performed by at least one processor, a) A step of receiving object information about a personalized object of a specific user included in a first image collected in a preset user environment, and performing labeling on the personalized first object of the specific user; b) a step of assigning a label identical to the label of the first object to a second object having a similarity greater than or equal to a preset similarity between the embedding vector of the first object for which labeling has been completed and the embedding vector of the second object included in the second image collected in a preset user environment; and c) a step of generating an artificial intelligence model trained to output a control signal according to a command when receiving a command for a personalized object of a specific user, wherein the first image and the second image are used as a learning data set, and a labeling process for the first object and the second object is included. A method for creating an artificial intelligence model that recognizes personalized objects.
2. In paragraph 1, The above object information is, A natural language instruction including ownership information for the first object; and Containing multi-angle images of the first object, A method for creating an artificial intelligence model that recognizes personalized objects.
3. In paragraph 1, Step b) above, Repeating the step of assigning the same label as the label of the first object to a second object having a similarity greater than or equal to a preset value with the first object until the proportion of objects to which the label is not assigned is less than 10%. A method for creating an artificial intelligence model that recognizes personalized objects.
4. In paragraph 1, The above similarity is, It is calculated based on the cosine similarity between the embedding vector of the first object for which the labeling has been completed and the embedding vector of the second object. A method for creating an artificial intelligence model that recognizes personalized objects.
5. Communication module; at least one processor; and A memory electrically connected to the processor and storing at least one code to be executed by the processor, The above memory, when executed through the above processor, The processor receives object information about a personalized object of a specific user included in a first image collected in a preset user environment, performs labeling on the personalized first object of the specific user, and assigns a label identical to the label of the first object to a second object having a preset similarity or higher to the first object based on a similarity between an embedding vector of the first object for which labeling has been completed and an embedding vector of a second object included in a second image collected in a preset user environment, and stores code that causes an artificial intelligence model trained to output a control signal according to the command when a command for the personalized object of the specific user is received. A device for generating an artificial intelligence model that recognizes personalized objects.
6. In paragraph 5, The above object information is, A natural language instruction including ownership information for the first object; and Containing multi-angle images of the first object, A device for generating an artificial intelligence model that recognizes personalized objects.
7. In paragraph 5, The above memory is, The processor further stores code that causes the processor to repeatedly assign the same label as the label of the first object to a second object having a similarity greater than or equal to a preset value with the first object until the proportion of objects to which the label is not assigned is less than 10%. A device for generating an artificial intelligence model that recognizes personalized objects.
8. In paragraph 5, The above similarity is, It is calculated based on the cosine similarity between the embedding vector of the first object for which the labeling has been completed and the embedding vector of the second object. A device for generating an artificial intelligence model that recognizes personalized objects.
9. In a robot control method using an artificial intelligence model, A step of receiving a command from a user regarding the user's object; A step of searching for an object identical to the user's object among objects in images collected in a preset user environment; and A step of controlling the robot based on a control signal output through an artificial intelligence model that recognizes a personalized object, A method for controlling a robot using an artificial intelligence model.
10. In paragraph 9, The artificial intelligence model that recognizes the above personalized object is An image collected in the above-described user environment and an image of the user's object are used as a learning data set, and a labeling process for the object in the image collected in the above-described user environment and the user's object is included, and when a command for the user's object is received, the system is trained to output a control signal according to the command. A method for controlling a robot using an artificial intelligence model.
11. In paragraph 9, The step of receiving a command from the user for the user's object is: A step of inferring bounding box coordinates for the user's object based on the above command, The above command is, Contains ownership information about the object of the above user, A method for controlling a robot using an artificial intelligence model.
12. In paragraph 9, The step of controlling the robot based on a control signal output through an artificial intelligence model that recognizes the personalized object is as follows: The robot detects the user's object, obtains the three-dimensional coordinates of the user's object through a preset algorithm, and manipulates the object. A method for controlling a robot using an artificial intelligence model.
13. In a robot control system using an artificial intelligence model, Communication module; at least one processor; and A memory electrically connected to the processor and storing at least one code to be executed by the processor, The above memory, when executed through the above processor, The processor stores a code that causes the robot to receive a command from the user about the user's object, search for an object identical to the user's object among objects collected in an image in a preset user environment, and control the robot based on a control signal output through an artificial intelligence model that recognizes the personalized object. Robot control system using artificial intelligence model.
14. In paragraph 13, The artificial intelligence model that recognizes the above personalized object is An image collected in the above-described user environment and an image of the user's object are used as a learning data set, and a labeling process for the object in the image collected in the above-described user environment and the user's object is included, and when a command for the user's object is received, the system is trained to output a control signal according to the command. Robot control system using artificial intelligence model.
15. In paragraph 13, The above memory is, The processor further stores code that causes the processor to receive a natural language instruction for the user's object and to infer bounding box coordinates for the user's object based on the natural language instruction, The above natural language instructions are, Contains ownership information about the object of the above user, Robot control system using artificial intelligence model.
16. In paragraph 13, The above memory is, The processor further stores a code that causes the robot to detect the user's object and manipulate the object by obtaining three-dimensional coordinates of the user's object through a preset algorithm. Robot control system using artificial intelligence model.
Citation Information
Patent Citations
Method of measuring adhesiveness for secondary battery
KR1020250147452A
Video monitoring system and installation method thereof
KR102211007B1
Method of detecting omitted or wrong result of bounding box work and computer apparatus conducting thereof
KR102337693B1
A method for generating a training dataset
KR102491025B1
Update of local features model based on correction to robot action
KR102517457B1