Method and apparatus for generating natural language instruction on basis of visual recognition

A visual recognition-based system addresses inefficiencies in manual input by automatically detecting and generating natural language instructions for objects, improving robot manipulation accuracy and adaptability.

WO2025178174A1PCT designated stage Publication Date: 2025-08-28SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/006527
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2024-05-14
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing sensor systems require manual input of object coordinates and commands, are inefficient for recognizing objects of varying sizes or multiple objects, and lack the ability to generate accurate natural language instructions based on visual recognition.

Method used

A method and device for visually detecting and extracting object features, generating natural language instructions based on criteria and requirements, using a visual recognition-based system that includes a communication module, processor, and memory to recognize objects, extract features, and generate instructions.

Benefits of technology

Enables rapid adaptation and improved instruction generation capabilities, accurately extracting features and generating natural language instructions for multiple objects, even in environments with limited data, enhancing robot manipulation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024006527_28082025_PF_FP_ABST
    Figure KR2024006527_28082025_PF_FP_ABST
Patent Text Reader

Abstract

One embodiment of the present disclosure provides a method for generating a natural language instruction on the basis of visual recognition. This method comprises the steps of: recognizing one or more objects in an image from the image; extracting features of the objects; and generating a natural language instruction for the objects according to pre-configured criteria and requirements on the basis of the features.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for generating natural language instructions based on visual recognition

[0001] The present invention relates to a method and device for generating natural language instructions based on visual recognition, and more particularly, to a technology for visually detecting and extracting information on the location, properties, and categories of an object, and generating natural language instructions composed in human language based on the information according to standards and requirements.

[0002] Recognizing objects and changing their positions in a space consisting of an image or a certain area requires precise technology.

[0003] Previously, the accuracy of sensor object recognition was low, requiring administrators to manually input object coordinates and the corresponding actions in the form of commands. However, this method required objects to remain in a fixed location, and entering commands for objects of various sizes or multiple objects was time-consuming.

[0004] Accordingly, the need for technology that allows sensors to recognize images or objects located in a certain space, extract features of the objects, and generate natural language instructions based on criteria or requirements between objects is increasing.

[0005] The present invention is intended to solve the problems of the prior art described above, and aims to visually detect and extract information on the location, properties and categories of an object, and based on this, generate natural language instructions composed of human language according to criteria and requirements.

[0006] However, the technical task that this embodiment seeks to achieve is not limited to the technical task described above, and other technical tasks may exist.

[0007] As a technical means for achieving the above-described technical task, an embodiment according to the first aspect of the present disclosure provides a method for generating natural language instructions based on visual recognition. The method comprises the steps of recognizing at least one object within an image from the image, extracting features of the objects, and generating natural language instructions for the objects based on the features and according to preset criteria and requirements.

[0008] In addition, an embodiment according to a second aspect of the present disclosure provides a visual recognition-based natural language instruction generation device. The device includes a communication module, at least one processor, and a memory electrically connected to the processor and storing at least one code to be executed by the processor, wherein the memory stores a code that, when executed through the processor, causes the processor to recognize at least one object within the image from an image, extract features of the objects, and generate a natural language instruction for the object based on preset criteria and requirements based on the features.

[0009] The present invention enables rapid adaptation even in situations where image data is small, and can improve instruction generation capabilities as learning data accumulates.

[0010] The present invention can accurately extract features of objects under the same conditions and generate natural language instructions according to relationship values ​​for multiple objects.

[0011] FIG. 1 is a drawing illustrating a server and a terminal and robot connected to the server in communication with one embodiment of the present invention.

[0012] Figure 2 is a drawing showing the detailed configuration of the server illustrated in Figure 1.

[0013] Figures 3 to 10 are exemplary diagrams showing the process of a robot control device according to one embodiment of the present invention.

[0014] Figure 11 is an exemplary diagram showing different image inference methods according to one embodiment of the present invention.

[0015] Figures 12 and 14 are tables comparing inference results of an image inference framework according to one embodiment of the present invention.

[0016] FIG. 15 is a flowchart illustrating the sequence of a visual recognition-based natural language instruction generation method according to another embodiment of the present invention.

[0017] FIG. 16 is a flowchart illustrating the sequence of an artificial intelligence model creation method using a visual recognition-based natural language instruction creation method according to another embodiment of the present invention.

[0018] Figure 17 is a flowchart illustrating the sequence of a robot control method to which an artificial intelligence model for generating a robot control signal is applied according to another embodiment of the present invention.

[0019] Hereinafter, the present disclosure will be described in detail with reference to the attached drawings. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, the attached drawings are only intended to facilitate understanding of the embodiments disclosed in the present specification, and the technical concepts disclosed in the present specification are not limited by the attached drawings. All terms, including technical and scientific terms, used herein should be interpreted as having meanings generally understood by a person of ordinary skill in the technical field to which the present disclosure pertains. Terms defined in the dictionary should be interpreted as having additional meanings consistent with the relevant technical literature and the present disclosure, and shall not be interpreted in an extremely ideal or restrictive sense unless otherwise defined.

[0020] In order to clearly explain the present invention in the drawings, parts irrelevant to the description have been omitted, and the size, shape, and appearance of each component shown in the drawings may be modified in various ways. Identical / similar parts throughout the specification are given identical / similar drawing reference numerals.

[0021] Throughout the specification, when a part is said to be "connected (connected, in contact with, or coupled)" to another part, this includes not only cases where it is "directly connected (connected, in contact with, or coupled)" but also cases where it is "indirectly connected (connected, in contact with, or coupled)" with another member in between. Furthermore, when a part is said to "include (have or provide)" a certain component, this does not mean that it excludes other components, but rather that it may "include (have or provide)" other components, unless otherwise specifically stated.

[0022] In this specification, the term 'unit' includes a unit realized by hardware, a unit realized by software, and a unit realized using both. In addition, one unit may be realized by using two or more pieces of hardware, and two or more units may be realized by one piece of hardware. Meanwhile, the '~ unit' is not limited to software or hardware, and the '~ unit' may be configured to be in an addressable storage medium or may be configured to reproduce one or more processors. Therefore, as an example, the '~ unit' includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and '~sub-units' may be combined into a smaller number of components and '~sub-units' or further separated into additional components and '~sub-units'. Furthermore, the components and '~sub-units' may be implemented to activate one or more CPUs within the device or secure multimedia card.

[0023] The suffixes "module" and "part" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. Furthermore, in describing the embodiments disclosed herein, detailed descriptions of related known technologies have been omitted if they are deemed to obscure the gist of the embodiments disclosed herein.

[0024] As used herein, ordinal terms such as "first," "second," etc., are used solely to distinguish one component from another and do not limit the order or relationship of the components. For example, the first component of the present disclosure may be referred to as the "second component," and similarly, the second component may also be referred to as the "first component." As used herein, singular forms should be construed to include plural forms, unless explicitly stated otherwise.

[0025] The "user terminal" mentioned below may be implemented as a computer or portable terminal that can access a server or other terminal via a network. Here, the computer may include, for example, a notebook, desktop, laptop, VR HMD (e.g., HTC VIVE, Oculus Rift, GearVR, DayDream, PSVR, etc.) equipped with a web browser. Here, the VR HMD includes all of the stand-alone models implemented independently for PC (e.g., HTC VIVE, Oculus Rift, FOVE, Deepon, etc.), mobile (e.g., GearVR, DayDream, Storm Magic, Google Cardboard, etc.), and console (PSVR) (e.g., Deepon, PICO, etc.). A portable terminal is, for example, a wireless communication device that ensures portability and mobility, and may include not only a smart phone, a tablet PC, and a wearable device, but also various devices equipped with communication modules such as Bluetooth (BLE, Bluetooth Low Energy), NFC, RFID, ultrasonic, infrared, WiFi, and LiFi. In addition, a "network" refers to a connection structure that enables information exchange between each node, such as terminals and servers, and includes a local area network (LAN), a wide area network (WAN), the Internet (WWW: World Wide Web), wired and wireless data communication networks, telephone networks, and wired and wireless television communication networks.Examples of wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, Bluetooth, infrared, ultrasonic, visible light communication (VLC), and LiFi.

[0026] FIG. 1 is a drawing illustrating a server and a terminal and robot connected to the server in communication with one embodiment of the present invention.

[0027] Referring to FIG. 1, in one example, the server (100) may be at least one of a natural language instruction generation device, an artificial intelligence model generation device, and a robot driving device.

[0028] The natural language instruction generation device can be connected to a terminal (200).

[0029] The natural language instruction generation device may be implemented as a cloud computing server, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). Furthermore, the natural language instruction device may be implemented in a private cloud, public cloud, or hybrid cloud system, but the scope of the present invention is not limited thereto.

[0030] A natural language instruction generation device receives an image from a terminal (200). However, the image reception is not limited to the terminal (200), and the image may be received from at least one of a storage device storing at least one image or a photographing device including a camera connected to a network. The natural language instruction generation device recognizes at least one object (subject) in the image received from the terminal (200) and can extract features of the objects. Then, based on the features, the device can generate a natural language instruction corresponding to the object according to preset criteria and conditions.

[0031] A terminal (200) that is connected to a natural language instruction generation device can transmit an image to the salmon instruction device and input a preset natural language instruction to the salmon instruction device.

[0032] The artificial intelligence model generation device can be connected to a terminal (200).

[0033] The AI ​​model generation device may be configured as a cloud computing server, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). Furthermore, the AI ​​model generation device may be configured as a private cloud, public cloud, or hybrid cloud system, but the scope of the present invention is not limited thereto.

[0034] An artificial intelligence model generation device receives an image from a terminal (200). However, the image reception is not limited to the terminal (200), and may receive an image from at least one of a storage device storing at least one image or a photographing device including a camera connected to a network. The artificial intelligence model generation device may recognize at least one object (object) in the image received from the terminal (200) and extract features of the objects. Then, based on the features, the device may generate natural language instructions corresponding to the objects according to preset criteria and conditions. Then, an artificial intelligence model may be generated that outputs a driving signal of the robot (300) using a learning data set including images, features, and natural language instructions as input values. In this case, the artificial intelligence model may be a visual grounding (VG) model.

[0035] A terminal (200) connected to an artificial intelligence model generation device can transmit an image to the artificial intelligence model generation device and input a preset natural language instruction to the artificial intelligence model generation device.

[0036] The robot driving device can be connected to the terminal (200) and the robot (300).

[0037] The robot drive unit may be configured as a cloud computing server, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). Furthermore, the robot drive unit may be configured as a private cloud, public cloud, or hybrid cloud system, but the scope of the present invention is not limited thereto.

[0038] The robot driving device receives an image from the terminal (200). However, the image reception is not limited to the terminal (200), and the image may be received from at least one of a storage device storing at least one image or a photographing device including a camera connected to a network. The robot driving device can recognize at least one object (subject) in the image received from the terminal (200) and extract the characteristics of the objects.

[0039] And, based on the features, natural language instructions corresponding to the object can be generated according to preset criteria and conditions. And, an artificial intelligence model can be generated that outputs a driving signal of the robot (300) by using a learning data set including images, features, and natural language instructions as input values. At this time, the artificial intelligence model may be a visual grounding model (Visual Grounding (VG) Model). And, the robot (300) can be driven through the artificial intelligence model.

[0040] A terminal (200) connected to a robot driving device can transmit an image to the robot driving device and input a preset natural language instruction to the robot driving device.

[0041] A robot (300) connected to a robot driving device and communicating with the robot driving device can receive a driving signal corresponding to a natural language instruction input by a terminal (200) from the robot driving device. In response to the received driving signal, the robot can detect the position of an object, obtain the three-dimensional coordinates of the object through a preset algorithm, and manipulate the object by predicting the target position.

[0042] Figure 2 is a drawing showing the detailed configuration of the server illustrated in Figure 1.

[0043] Referring to FIG. 2, in one example, the server (100) may be at least one of a natural language instruction generation device, an artificial intelligence model generation device, and a robot driving device. In this case, the natural language instruction generation device, the artificial intelligence model generation device, and the robot driving device may be referred to as a natural language instruction generation system, an artificial intelligence model generation system, and a robot driving system, respectively.

[0044] The natural language instruction generation device may include a communication module (110), a processor (120), and a memory (130).

[0045] In a natural language instruction generation device, a communication module (110) may include a device including hardware and software necessary to transmit and receive signals such as control signals or data signals through a wired or wireless connection with another network device.

[0046] In a natural language instruction generation device, the communication module (110) can receive an image from the terminal (200). However, this is not limited thereto, and if necessary, the communication module (110) can receive an image from a storage device in which an image is stored or an external device that already includes a photographing unit (camera) other than the terminal (200). In addition, the communication module (110) can receive a natural language instruction from the terminal.

[0047] In a natural language instruction generation device, the processor (120) may include various types of devices that control and process data. In a natural language instruction generation device, the processor (120) may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program.

[0048] In the natural language instruction generation device, the processor (120) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.

[0049] In a natural language instruction generation device, the processor (120) performs an operation according to a code stored in a memory (130).

[0050] In a natural language instruction generation device, the memory (130) can store at least one of information and data input to the communication module (110), information and data required for a function performed by the processor (120), and data generated according to the execution of the processor (120).

[0051] In the natural language instruction generation device, the memory (130) should be interpreted as a general term for a non-volatile storage device that maintains stored information even when no power is supplied and a volatile storage device that requires power to maintain the stored information. The memory (130) may include a magnetic storage media or a flash storage media in addition to a volatile storage device that requires power to maintain the stored information, but the scope of the present invention is not limited thereto.

[0052] In a natural language instruction generation device, a memory (130) is electrically connected to a processor (120) and stores at least one code to be executed by the processor (120). The memory (130) stores a code that, when executed by the processor (120), causes the processor (120) to perform the following functions and procedures.

[0053] In a natural language instruction generation device, a code that causes at least one object within an image to be recognized is stored in the memory (130).

[0054] In a natural language instruction generation device, a code that causes the extraction of features of objects is stored in the memory (130). For example, the features may include location information, attribute information, and category information of the object.

[0055] In a natural language instruction generation device, a code that causes a natural language instruction to be generated for an object according to preset criteria and requirements based on characteristics is stored in the memory (130). For example, a code that causes a template to be generated by combining position information, attribute information, and category information according to a preset algorithm and to generate a natural language instruction according to the template may be stored. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object. At least one or more of the position information, attribute information, and category information and the relationship value are combined to form a template, and a natural language instruction for an operation performed by the robot may be generated based on the template. In this case, the relationship value may include at least one of the relative position and final position information of a plurality of objects.

[0056] The artificial intelligence model generation device may include a communication module (110), a processor (120), and a memory (130).

[0057] In the artificial intelligence model generation device, the communication module (110) may include a device including hardware and software necessary to transmit and receive signals such as control signals or data signals through a wired or wireless connection with another network device.

[0058] In the artificial intelligence model generation device, the communication module (110) can receive images from the terminal (200). However, this is not limited thereto, and if necessary, the communication module (110) can receive images from a storage device storing images or an external device including a camera, other than the terminal (200). In addition, the communication module (110) can receive natural language instructions from the terminal.

[0059] In an artificial intelligence model generation device, the processor (120) may include various types of devices that control and process data. In an artificial intelligence model generation device, the processor (120) may refer to a data processing device built into hardware that has a physically structured circuit to perform functions expressed by code or commands included in a program.

[0060] In the artificial intelligence model generation device, the processor (120) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.

[0061] In an artificial intelligence model generation device, the processor (120) performs operations according to code stored in the memory (130).

[0062] In the artificial intelligence model generation device, the memory (130) can store at least one of information and data input to the communication module (110), information and data required for a function performed by the processor (120), and data generated according to the execution of the processor (120).

[0063] In the artificial intelligence model generation device, the memory (130) should be interpreted as a general term for a non-volatile storage device that maintains stored information even when no power is supplied, and a volatile storage device that requires power to maintain the stored information. The memory (130) may include a magnetic storage media or a flash storage media in addition to a volatile storage device that requires power to maintain the stored information, but the scope of the present invention is not limited thereto.

[0064] In the artificial intelligence model generation device, the memory (130) is electrically connected to the processor (120) and stores at least one code that is executed by the processor (120). The memory (130) stores a code that, when executed by the processor (120), causes the processor (120) to perform the following functions and procedures.

[0065] In the artificial intelligence model generation device, a code that causes at least one object within an image to be recognized is stored in the memory (130).

[0066] In the artificial intelligence model generation device, a code that causes the extraction of features of objects is stored in the memory (130). For example, the features may include location information, attribute information, and category information of the object.

[0067] In the artificial intelligence model generation device, a code that causes a natural language instruction to be generated for an object according to preset criteria and requirements based on features is stored in the memory (130). For example, a code that causes a template to be generated by combining position information, attribute information, and category information according to relationship values ​​generated based on a preset algorithm, and a code that causes a natural language instruction corresponding to the template to be generated may be stored. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object, and at least one or more of the position information, attribute information, and category information is combined with relationship values ​​to form a template, and a natural language instruction for an operation to be performed by the robot can be generated based on the template.

[0068] In the artificial intelligence model generation device, a code that causes an artificial intelligence model trained to output a control signal for a robot based on a learning data set including images, features, and natural language instructions is stored in the memory (130). For example, the generated artificial intelligence model may be trained to output a control signal for a robot based on a learning data set including images, features, and natural language instructions, based on a Visual Grounding model trained to display a recognition target object as a bounding box within an image.

[0069] The robot driving device may include a communication module (110), a processor (120), and a memory (130).

[0070] In a robot driving device, a communication module (110) may include a device including hardware and software necessary to transmit and receive signals such as control signals or data signals through a wired or wireless connection with another network device.

[0071] In the robot driving device, the communication module (110) can receive an image from the terminal (200). However, this is not limited thereto, and if necessary, the communication module (110) can receive an image from a storage device in which an image is stored or an external device that already includes a photographing unit (camera) other than the terminal (200). In addition, the communication module (110) can receive a natural language instruction from the terminal.

[0072] In a robot driving device, the processor (120) may include various types of devices that control and process data. In a robot driving device, the processor (120) may refer to a data processing device built into hardware that has a physically structured circuit to perform a function expressed by a code or command included in a program.

[0073] In the robot driving device, the processor (120) may be implemented in the form of a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.

[0074] In a robot driving device, the processor (120) performs operations according to the code stored in the memory (130).

[0075] In the robot driving device, the memory (130) can store at least one of information and data input to the communication module (110), information and data required for functions performed by the processor (120), and data generated according to the execution of the processor (120).

[0076] In a robot driving device, the memory (130) should be interpreted as a general term for a non-volatile storage device that maintains stored information even when no power is supplied and a volatile storage device that requires power to maintain the stored information. The memory (130) may include a magnetic storage media or a flash storage media in addition to a volatile storage device that requires power to maintain the stored information, but the scope of the present invention is not limited thereto.

[0077] In the robot driving device, the memory (130) is electrically connected to the processor (120) and stores at least one code that is executed by the processor (120). The memory (130) stores a code that, when executed by the processor (120), causes the processor (120) to perform the following functions and procedures.

[0078] In the robot driving device, a code that causes the memory (130) to recognize at least one object within an image is stored.

[0079] In the robot driving device, a code that causes the extraction of features of objects is stored in the memory (130). For example, the features may include location information, attribute information, and category information of the object.

[0080] In the robot driving device, a code that causes a natural language instruction to be generated for an object according to preset criteria and requirements based on characteristics is stored in the memory (130). For example, a code that causes a template to be generated by combining position information, attribute information, and category information according to relationship values ​​generated based on a preset algorithm, and a code that causes a natural language instruction corresponding to the template to be generated may be stored. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object, and at least one of the position information, attribute information, and category information is combined with relationship values ​​to form a template, and a natural language instruction for an operation performed by the robot can be generated based on the template.

[0081] In a robot driving device, a memory (130) stores code that causes an artificial intelligence model trained to output a control signal for the robot based on a learning data set including images, features, and natural language instructions. For example, the generated artificial intelligence model may be trained to output a control signal for the robot based on a learning data set including images, features, and natural language instructions, based on a Visual Grounding model trained to display a recognition target object as a bounding box within an image.

[0082] In a robot driving device, a memory (130) stores code that causes the robot to operate based on a control signal output through an artificial intelligence model. For example, when a natural language instruction is input, the robot can operate by predicting the location and target location of a target object among objects through an artificial intelligence model. In addition, the robot can detect an object, obtain the object's three-dimensional coordinates through a preset algorithm, and manipulate the object.

[0083] Figures 3 to 10 are exemplary diagrams showing the process of a robot control device according to one embodiment of the present invention.

[0084] Referring to Figure 3, conventional visual recognition techniques train a visual grounding model using a public data set (410). Assuming that the public data set (410) is a good example of an object, a frozen model is used to infer the system's progress. However, most public data sets are very different from the environments in which a robot operates. In the present invention, a new data set (420) is used to address the gap between the public data set (410), which can limit the robot's operational capabilities, and data observed in the robot's environment.

[0085] Before explaining FIG. 4, the server (100) illustrated in FIG. 2 can be operated based on a framework called GVCCI (Grounding Vision to Ceaselessly Created Instructions).

[0086] Referring to FIG. 4, a first framework (510) applied to a server (100) according to one embodiment of the present invention automatically generates natural language instructions through an original image in an operation area of ​​grounding vision for continuously generated natural language instructions, and accordingly, can adjust a visual grounding model as shown in 520 of FIG. 4.

[0087] At this time, the visual feature extraction module, instruction generation module, visual grounding model, and manipulation module that constitute the framework may be included in the processor of the server (100). In addition, the probabilistic volatile buffer may be included in the memory of the server (100).

[0088] To elaborate, the visual feature extraction module extracts the category, location, and properties of objects from an image. The instruction generation module generates realistic pick-up and drop instructions based on visual information. The probabilistic volatile buffer stores image, instruction, and bounding box features and is a memory that probabilistically forgets data. The visual grounding model infers the target object and target location using visual information and instructions. The manipulation module plans the manipulation trajectory of the robotic arm based on information obtained from the visual grounding model.

[0089] Referring to 620 and 630 shown in FIG. 5, the second framework (610) applied to the server (100) according to one embodiment of the present invention can be easily utilized in all areas where the robot operates because it supports the robot to adapt to a specific operating environment using only the original image.

[0090] Referring to FIG. 6, the entire framework applied to the server can be confirmed, and referring to 710 of FIG. 6, the server (100) can enable the robot to continuously learn visual grounding without the supervision of a worker, and referring to 720 of FIG. 6, the robot can perform LGRM (Language-Guided Robotic Manipulation) to make inference.

[0091] Referring to FIGS. 7 to 9, the robot described with reference to 710 of FIG. 6 learns visual grounding by first extracting an object and its corresponding features from an original image (810), as shown in FIG. 7 at 820. Next, as shown in FIG. 8 at 920, the features of the extracted object are used to generate a natural language construction through a predefined template (910). At this time, in the template (910), a, c, and b are features for attribute information, category information, and location information for the detected object, respectively, A and C are features for attribute information and category information for objects related to the detected object, and R represents a relationship value between the detected object and its related object. At this time, the attribute information is information indicating the material or color of the object, the category information is information indicating the type or category of the object, and the location information is information indicating the coordinates of the object. These characteristics can be extracted through a heuristic algorithm.

[0092] For example, referring to the template (910) of FIG. 8, {c}+{R}+{A}+{C} as in Ⅳ can generate a natural language construction as "can next to blue box" as an example. This is a natural language construction that combines "can", which is the category information (c) of the detected object, "blue box", which is the attribute information (A) and category information (C) of the object related to the detected object, and "next", which is the relationship value (R), and has the meaning of "can next to the blue box."

[0093] And {a}+{c}+{R}+{A}+{C} can generate a natural language construction as "yellow can on the right of blue box" as an example. This is a natural language construction that combines "yellow can", which combines the attribute information (a) and category information (c) of the detected object, "blue box", which combines the attribute information (A) and category information (C) of objects related to the detected object, and "on the right of", which is the relation value (R), to mean "yellow can on the right of the blue box."

[0094] Next, a triplet (1010) consisting of the original image, the coordinates of the target object (object), and a natural language instruction is stored in a buffer (1020). At this time, the triplet (1010) stored in the buffer may be deleted starting from the oldest. This is because data cannot be stored infinitely in the buffer. In addition, the server (100) can update the visual grounding model using the stored data.

[0095] Referring to FIG. 10, when a worker inputs a natural language instruction (1110) to a robot to manipulate an object, the position to which the object will be moved can be predicted through a visual grounding model (1120). Then, the RANSAC algorithm outputs the 3D position of the object in a 3D space (1130), and the robot can be made to move the "yellow cup" as in the natural language instruction (1110).

[0096] Figure 11 is an exemplary diagram showing different image inference methods according to one embodiment of the present invention.

[0097] Referring to FIG. 11, the Zero-shot Performance (1210), which is a conventional visual recognition technology, the application of PseudoQ to the framework of the present invention (1220), and the GVCCI framework (1230) operating the server (100) of the present invention can be compared and confirmed.

[0098] Figures 12 to 14 are tables comparing inference results of an image inference framework according to one embodiment of the present invention.

[0099] Referring to FIGS. 12 and 13, the accuracy (1310) when the data set is "0", the accuracy (1320) when PseudoQ is applied, and the accuracy (1410) of the framework of the present invention can be compared. The characteristic of this table is that in the case of PseudoQ, the performance is the best when the data set is 33 sheets, and then the performance deteriorates as the data set increases, whereas in the case of the framework of the present invention, the performance increases as the data set increases.

[0100] To explain this in detail, as shown in the results of FIGS. 12 and 13, the GVCCI operating the server (100) of the present invention was found to outperform the existing methods in various test sets (Test-H, Test-R, Test-E) and two backbone models (MDETR, OFA). In particular, the GVCCI operating the server (100) of the present invention showed excellent performance improvement through rapid adaptation even in situations where there were few observed scenes in the environment (8, 33), which shows that the GVCCI operating the server (100) of the present invention has the ability to continuously improve performance as learning data accumulates. On the other hand, the PseudoQ method showed minimal performance improvement due to overfitting in the early stages.

[0101] And as shown in Fig. 14, in an online experiment, the GVCCI operating the server (100) of the present invention showed high performance in actual robot operation. In four criteria of pick inference, pick operation, place inference, and overall pick-and-place accuracy, the GVCCI operating the server (100) of the present invention surpassed existing methods, and in particular, the Zero-Shot method had a low success rate in pick operation, which caused a problem in that it could not effectively pick up actual objects.

[0102] On the other hand, the robot using the GVCCI operating the server (100) of the present invention was found to be relatively good at picking up objects, showing significantly improved performance. The advantage of the improved performance in the pick operation is that the GVCCI operating the server (100) of the present invention more accurately detects the coordinates of objects when learning through objects in the environment, thereby increasing the manipulation success rate. These experimental results confirm that the present invention provides superior performance and adaptability compared to existing technologies in the field of language-guided robot manipulation.

[0103] FIG. 15 is a flowchart illustrating the sequence of a visual recognition-based natural language instruction generation method according to another embodiment of the present invention.

[0104] The visual recognition-based natural language instruction generation method described below can be performed by the visual recognition-based natural language instruction generation device described above with reference to FIGS. 1 to 14. Therefore, the contents of the embodiments of the present disclosure described above with reference to FIGS. 1 to 14 can be equally applied to the embodiments described below, and any overlapping contents with the above description will be omitted below. The steps described below do not necessarily have to be performed in order, the order of the steps can be set in various ways, and the steps can be performed almost simultaneously.

[0105] Referring to FIG. 15, a method for generating natural language instructions based on visual recognition includes an object recognition step (S110), a feature extraction step (S120), and a natural language instruction generation step (S130).

[0106] The object recognition step (S110) is a step of recognizing at least one object within an image from an image.

[0107] The feature extraction step (S120) is a step for extracting features of objects. For example, the features may include location information, attribute information, and category information of the object.

[0108] The natural language instruction generation step (S130) is a step of generating a natural language instruction for an object according to preset criteria and requirements based on features. For example, a template can be generated by combining position information, attribute information, and category information according to relationship values ​​generated based on a preset algorithm, and a natural language instruction corresponding to the template can be generated. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object. At least one or more of the position information, attribute information, and category information is combined with relationship values ​​to form a template, and a natural language instruction for an operation performed by the robot can be generated based on the template. At this time, the relationship value can include at least one of the relative positions and final position information of a plurality of objects.

[0109] FIG. 16 is a flowchart illustrating the sequence of an artificial intelligence model creation method using a visual recognition-based natural language instruction creation method according to another embodiment of the present invention.

[0110] The artificial intelligence model generation method using the natural language instruction generation method described below can be performed by the artificial intelligence model generation device using the natural language instruction generation method described above with reference to FIGS. 1 to 14. Therefore, the contents of the embodiments of the present disclosure described above with reference to FIGS. 1 to 14 can be equally applied to the embodiments described below, and any content overlapping with the above description will be omitted below. The steps described below do not necessarily have to be performed in order, the order of the steps can be set in various ways, and the steps can be performed almost simultaneously.

[0111] Referring to FIG. 16, the method for generating an artificial intelligence model using a visual recognition-based natural language instruction generation method includes an object recognition step (S210), a feature extraction step (S220), a natural language instruction generation step (S230), and an artificial intelligence model generation step (S240).

[0112] The object recognition step (S210) is a step of recognizing at least one object within an image from an image.

[0113] The feature extraction step (S220) is a step for extracting features of objects. For example, the features may include object location information, attribute information, and category information.

[0114] The natural language instruction generation step (S230) is a step of generating a natural language instruction for an object according to preset criteria and requirements based on features. For example, a template can be generated by combining position information, attribute information, and category information according to relationship values ​​generated based on a preset algorithm, and a natural language instruction corresponding to the template can be generated. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object. At least one or more of the position information, attribute information, and category information is combined with relationship values ​​to form a template, and a natural language instruction for an operation performed by the robot can be generated based on the template. At this time, the relationship value can include at least one of the relative positions and final position information of a plurality of objects.

[0115] The AI ​​model generation step (S240) is a step for generating an AI model trained to output control signals for a robot based on a learning data set including images, features, and natural language instructions. For example, the generated AI model may be trained to output control signals for a robot based on a learning data set including images, features, and natural language instructions, based on a Visual Grounding model trained to display a recognition target object as a bounding box within an image.

[0116] Figure 17 is a flowchart illustrating the sequence of a robot control method to which an artificial intelligence model for generating a robot control signal is applied according to another embodiment of the present invention.

[0117] The visual recognition-based natural language instruction generation method, the artificial intelligence model generation method using the same, and the robot driving method to which the artificial intelligence model is applied, which will be described below, can be performed by the visual recognition-based natural language instruction generation method, the artificial intelligence model generation method using the same, and the robot device to which the artificial intelligence model is applied, which have been described above with reference to FIGS. 1 to 14. Therefore, the contents of the embodiments of the present disclosure described above with reference to FIGS. 1 to 14 can be equally applied to the embodiments to be described below, and any contents overlapping with the above description will be omitted below. The steps described below do not necessarily have to be performed in order, the order of the steps can be set in various ways, and the steps can be performed almost simultaneously.

[0118] Referring to Fig. 17, a robot control method in which an artificial intelligence model for generating a robot control signal is applied includes a visual recognition-based natural language instruction generation method including an object recognition step (S310), a feature extraction step (S320), a natural language instruction generation step (S330), an artificial intelligence model generation step (S340), and a robot driving step (S350).

[0119] The object recognition step (S310) is a step of recognizing at least one object within an image from an image.

[0120] The feature extraction step (S320) is a step for extracting features of objects. For example, the features may include object location information, attribute information, and category information.

[0121] The natural language instruction generation step (S330) is a step of generating a natural language instruction for an object according to preset criteria and requirements based on features. For example, a template can be generated by combining position information, attribute information, and category information according to relationship values ​​generated based on a preset algorithm, and a natural language instruction corresponding to the template can be generated. In addition, the position information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the type or category of the object. At least one or more of the position information, attribute information, and category information is combined with relationship values ​​to form a template, and a natural language instruction for an operation performed by the robot can be generated based on the template. At this time, the relationship value can include at least one of the relative positions and final position information of a plurality of objects.

[0122] The AI ​​model generation step (S340) is a step for generating an AI model trained to output control signals for a robot based on a learning data set including images, features, and natural language instructions. For example, the generated AI model may be trained to output control signals for a robot based on a learning data set including images, features, and natural language instructions, based on a Visual Grounding model trained to display a recognition target object as a bounding box within an image.

[0123] The robot driving step (S350) is a step that drives the robot based on control signals output from an artificial intelligence model. For example, when natural language instructions are input, the robot can be driven by predicting the location of a target object and its target location among objects using the artificial intelligence model. Furthermore, the robot can detect the object, obtain its 3D coordinates using a preset algorithm, and manipulate the object.

[0124] An embodiment of the present invention may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include all computer storage media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.

[0125] Although the methods and systems of the present invention have been described with respect to specific embodiments, some or all of their components or operations may be implemented using a computer system having a general-purpose hardware architecture.

[0126] Those skilled in the art will appreciate that the present disclosure can be easily modified into other specific forms based on the above description without changing the technical spirit or essential characteristics of the present disclosure. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. The scope of the present disclosure is indicated by the following claims, and all changes or modifications derived from the meaning and scope of the claims and their equivalents should be construed as being included in the scope of the present disclosure. The scope of the present application is indicated by the following claims rather than the above detailed description, and all changes or modifications derived from the meaning and scope of the claims and their equivalents should be construed as being included in the scope of the present application.

Claims

1. A natural language instruction generation method based on visual recognition performed by a natural language instruction generation device, a) a step of recognizing at least one object within the image from the image; b) a step of extracting features of the above objects; and c) A method for generating natural language instructions based on visual recognition, comprising a step of generating natural language instructions for the object according to preset criteria and requirements based on the above characteristics.

2. In paragraph 1, In step b), the above features are A method for generating natural language instructions based on visual recognition, the method including location information, attribute information, and category information of the above object.

3. In paragraph 2, Step c) above, A template is created by combining the relationship values ​​generated based on a preset algorithm based on the above location information, attribute information, and category information, A method for generating natural language instructions based on visual recognition, wherein the natural language instructions are generated according to the above template.

4. In paragraph 3, The location information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the types or categories of the object. The template is configured by combining at least one of the location information, attribute information, and category information with the relationship value, A method for generating natural language instructions based on visual recognition, wherein natural language instructions for actions performed by a robot are generated based on the above template.

5. In paragraph 4, A method for generating natural language instructions based on visual recognition, wherein the above relationship value includes at least one of relative position information and final position information of a plurality of objects.

6. Communication module; at least one processor; and A memory electrically connected to the processor and storing at least one code to be executed by the processor, The above memory, when executed through the processor, causes the processor to: A visual recognition-based natural language instruction generation device storing code that causes the device to recognize at least one object in an image from an image, extract features of the objects, and generate natural language instructions for the objects based on the features according to preset criteria and requirements.

7. In paragraph 6, The above memory, when executed through the processor, causes the processor to: A visual recognition-based natural language instruction generation device, wherein the above features store code that causes the object to include location information, attribute information, and category information.

8. In paragraph 7, The above memory, when executed through the processor, causes the processor to: A template is created by combining the relationship values ​​generated based on a preset algorithm based on the above location information, attribute information, and category information, A visual recognition-based natural language instruction generation device storing code that causes the natural language instruction to be generated according to the above template.

9. In paragraph 8, The location information of the template is information indicating the coordinates of the object, the attribute information is information indicating at least one of the color or material of the object, and the category information is information indicating one of the types or categories of the object. The template is configured by combining at least one of the location information, attribute information, and category information with the relationship value, A visual recognition-based natural language instruction generation device that generates natural language instructions for actions performed by a robot based on the above template.

10. In paragraph 9, A visual recognition-based natural language instruction generation device, wherein the above relationship value includes at least one of relative position information and final position information of a plurality of objects.

Citation Information

Patent Citations

  • Method, apparatus, device and medium for generating captioning information of multimedia data

    KR102593440B1

  • Natural Language Based Computer Animation

    US20180293050A1

  • Self-Aware Visual-Textual Co-Grounded Navigation Agent

    US20200103911A1

  • Controlling interactive agents using multi-modal inputs

    WO2023104880A1