Method and system for learning visual-language-action model for physical ai

KR103013768B1Active Publication Date: 2026-09-02국립창원대학교산학협력단
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020250187938
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-09-02
Estimated Expiration
2045-12-02

Smart Images

  • Figure 112025135664859-PAT00001_ABST
    Figure 112025135664859-PAT00001_ABST
Patent Text Reader

Abstract

A method and system for training a Vision-Language-Action (VLA) model for physical AI are disclosed. A VLA model learning method according to one embodiment may include the steps of: generating visual embeddings by inputting visual information acquired from a robot work environment into a vision encoder; generating language embeddings by inputting at least one of natural language instructions and text input into a language encoder; generating state embeddings by inputting at least one of the robot's joint state, posture information, and motion history into a state encoder; inputting the visual embeddings, language embeddings, and state embeddings into a multimodal transformer model to learn the correlation between the embeddings and generate action tokens; constructing the learning dataset using a learning dataset composed of multiple robot work episodes, wherein the learning dataset is constructed using at least one of (i) a uniform random sampling method that randomly selects the location of a target object, and (ii) a hierarchical sampling method that divides the workspace into layers and collects episodes by layer; and obtaining a learned VLA model that predicts the robot's next action by updating the parameters of the multimodal transformer model to minimize the difference between the action tokens and the robot's actual motion data.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to the field of robot and physical AI technology, and more specifically, to a method and system for learning a Vision-Language-Action (VLA) model that predicts the next action by integrating visual information, language information, and state information of a robot, and furthermore, to a robot control method and a robot control system for generating and controlling the next action of a robot using the VLA model learned as described above. Background Technology

[0002] Recently, robot technology has expanded beyond manufacturing, logistics, and service sectors to include home environments, requiring them to perform a variety of tasks. These robot systems are primarily controlled by predefined programs or direct operator input; however, traditional methods such as coordinate-based control or teaching-based manipulation have limitations, as they rely heavily on expert know-how and lack adaptability to environmental changes.

[0003] Meanwhile, with the advancement of artificial intelligence technology, multimodal models that combine visual and linguistic information for understanding are emerging, and attempts to predict robot behavior by combining them with motion data are becoming increasingly active. However, conventional model-based robot control struggles to adequately reflect the diversity and complexity of real-world environments where visual, linguistic, and state information interact in a complex manner. Furthermore, limitations exist in terms of securing training data for stably generating the continuous actions required by the robot and in terms of the efficiency of the model structure.

[0004] Furthermore, while there is an increasing demand to flexibly control robots based on natural language commands, existing technologies suffer from problems such as information loss or inaccuracy during the process of converting complex linguistic instructions into actions that the robot can execute.

[0005] Under these circumstances, there is a growing need for model learning technologies that combine multimodal information to more accurately predict a robot's next action and reliably respond to various tasks in real-world environments.

[0006] [Prior Art No.]

[0007] Korean Patent Publication No. 10-2019-0057687 The problem to be solved

[0008] We provide a training method and system for a Vision-Language-Action (VLA) model for physical AI. means of solving the problem

[0009] A method for training a VLA model in a VLA model learning system that trains a Vision-Language-Action (VLA) model that predicts the next action of a robot by integrating visual information, language information, and state information of the robot, wherein the VLA model learning system is implemented by at least one computer device including at least one processor, and the VLA model learning method comprises: a step of generating a visual embedding by inputting visual information acquired from a robot work environment into a vision encoder by the at least one processor; a step of generating a language embedding by inputting at least one of natural language instruction and text input into a language encoder by the at least one processor; a step of generating a state embedding by inputting at least one of the robot's joint state, pose information, and motion history into a state encoder by the at least one processor; and a step of inputting the visual embedding, language embedding, and state embedding into a multimodal transformer model by the at least one processor to learn the correlation between the embeddings and generate an action token. A method for training a VLA model is provided, comprising the steps of: using a training dataset composed of multiple robot task episodes by the at least one processor, wherein the training dataset is configured using at least one of (i) a uniform random sampling method for randomly selecting the location of a target object, and (ii) a hierarchical sampling method for dividing a workspace into layers and collecting episodes by layer; and obtaining a trained VLA model that predicts the next action of a robot by updating the parameters of the multimodal transformer model to minimize the difference between the action token and the actual motion data of the robot by the at least one processor.

[0010] According to one aspect, the hierarchical sampling method may be characterized by being performed to reduce bias in the data distribution by dividing the workspace into multiple layers and collecting an equal number of episodes from each layer.

[0011] According to another aspect, the multimodal transformer model may be characterized by being configured to learn the correlation between the embeddings by performing self-attention on each of the visual embeddings, language embeddings, and state embeddings, and cross-attention between the embeddings.

[0012] According to another aspect, the parameter update of the multimodal transformer model may be characterized by being performed based on a loss function set to minimize the difference between the action token and the target action vector representing the robot's actual continuous motion.

[0013] According to another aspect, the vision encoder, language encoder, and state encoder are each fine-tuned based on a pre-trained model, and the fine-tuning may be characterized by being performed using visual information captured in a robot work environment and collected robot motion data.

[0014] According to another aspect, the step of generating the visual embedding may be characterized by extracting tokens in the form of image patches from input visual information and converting the tokens into embedding vectors using vision converter-based encoding.

[0015] According to another aspect, the step of generating the language embedding may be characterized by generating the language embedding using a language encoding that tokenizes natural language instructions and maps the tokens to the embedding space of a pre-trained language model.

[0016] According to another aspect, the step of generating the state embedding may be characterized by receiving at least one of the robot's joint angle, joint velocity, and end-effector pose as input, and generating the state embedding by projecting the value received as input into an embedding space using a state encoding neural network.

[0017] According to another aspect, the step of generating the action token may be characterized by generating the action token by reflecting the sequential meaning of the embeddings by further inputting a position encoding representing the temporal order of each embedding into a multimodal transformer model in addition to the visual embedding, language embedding, and state embedding.

[0018] According to another aspect, the parameter update of the multimodal transformer model may be characterized by being performed using a training epoch process in which a training dataset is input in batch units and the parameters are optimized based on the loss value for said batch.

[0019] A robot control method for generating the next action of a robot using a learned Vision-Language-Action (VLA) model, wherein the robot control method is executed by a computer device comprising at least one processor, and wherein the at least one processor inputs visual information acquired from a robot work environment into a vision encoder to generate a visual embedding; wherein the at least one processor inputs at least one of a user's natural language instruction and text input into a language encoder to generate a language embedding; wherein the at least one processor inputs at least one of the robot's joint state, posture information, and motion history into a state encoder to generate a state embedding; and wherein the at least one processor inputs the visual embedding, language embedding, and state embedding into the learned VLA model to generate an action token from the learned VLA model. The present invention provides a robot control method characterized by including the step of converting the action token into a robot driving command representing a joint command or end-effector trajectory of the robot by the at least one processor, and providing the robot driving command to a robot device to control the robot to perform the next action.

[0020] According to one aspect, the step of generating the language embedding may be characterized by converting a user's voice input into text using a speech recognition model, inputting the text into a large-scale language model to generate control instructions to be performed by a robot, and inputting the control instructions into the language encoder to generate the language embedding.

[0021] According to another aspect, the large-scale language model may be characterized by performing inference using a system prompt to convert natural language instructions into control commands in a format executable by the robot.

[0022] According to another aspect, the robot control method may further include the step of periodically collecting sensor feedback of the robot by the at least one processor, generating a new state embedding reflecting the sensor feedback, and then updating the robot's behavior in real time by repeatedly executing the learned VLA model using the new state embedding.

[0023] A computer program stored on a non-transient computer-readable recording medium is provided, which is combined with a computer device and executes the above method on the computer device.

[0024] A non-transient computer-readable recording medium is provided on which a computer program for executing the above method on a computer device is recorded.

[0025] A robot control system for generating the next action of a robot using a learned Vision-Language-Action (VLA) model comprises: a vision encoder configured to generate a vision embedding by receiving visual information acquired from a robot work environment as input; a language encoder configured to generate a language embedding by receiving at least one of a user's natural language instruction and text input as input; a state encoder configured to generate a state embedding by receiving at least one of the robot's joint state, posture information, and motion history as input; a multimodal inference module configured to generate an action token from a learned VLA model by receiving the vision embedding, language embedding, and state embedding as input; and a command converter configured to convert the action token into a robot driving command representing a robot's joint command or end-effector trajectory, wherein the robot is configured to perform the next action by receiving and executing the robot driving command generated by the command converter. Effects of the invention

[0026] We can provide a training method and system for a Vision-Language-Action (VLA) model for physical AI.

[0027] A robot control method and system for generating and controlling the next action of a robot using a learned VLA model can be provided.

[0028] According to embodiments of the present invention, by integrally utilizing visual information, language information, and robot state information, the next action of a robot can be predicted more accurately and consistently, thereby enabling the implementation of a physical AI system that performs operations stably even in various work environments.

[0029] In addition, by performing learning using multimodal information, the problems of environment dependency and lack of adaptability inherent in existing coordinate-based or teaching-based control methods can be resolved, and the reliability and precision of the process of converting complex natural language commands into control commands that the robot can execute can be improved.

[0030] Furthermore, the learning method according to the embodiments of the present invention improves the structuring and efficiency of data collection, thereby enabling the construction of a behavioral model with excellent generalization performance for various states, scenes, and tasks.

[0031] Furthermore, since it can be applied to lightweight model structures, it facilitates learning and inference even with limited computing resources, thereby having a positive effect of expanding the scope of experimental and industrial applications. Brief explanation of the drawing

[0032] FIG. 1 is a diagram illustrating an example of an overview of a full Vision-Language-Action (VLA) based system according to one embodiment of the present invention. FIG. 2 is a block diagram illustrating an example of the internal configuration of a VLA model learning system according to an embodiment of the present invention. FIG. 3 is a flowchart illustrating an example of a VLA model learning method according to an embodiment of the present invention. FIG. 4 is a block diagram illustrating an example of the internal configuration of a robot control system according to an embodiment of the present invention. FIG. 5 is a flowchart illustrating an example of a robot control method according to an embodiment of the present invention. FIG. 6 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. FIG. 7 is a diagram schematically illustrating the input processing and behavior generation processes of a visual-language-behavior model according to an embodiment of the present invention. FIG. 8 is a diagram illustrating an example of the overall processing flow in which a continuous control command of a robot is generated from a voice-based natural language input through a multimodal visual-language-action (VLA) model according to one embodiment of the present invention. FIG. 9 is a drawing illustrating screens showing an actual implementation example of a robot control system according to an embodiment of the present invention. FIG. 10 is a diagram illustrating an example of the overall processing flow in which a voice-based natural language command is converted into an executable control command of a robot in a robot control system according to an embodiment of the present invention. FIG. 11 is a diagram illustrating an example of a process in which a text-based command is converted into a robot control command in a robot control system according to an embodiment of the present invention. FIG. 12 is a diagram illustrating an example of the internal information processing flow of a visual-language-action model that generates continuous actions based on visual, language, and state information in a robot control system according to an embodiment of the present invention. FIG. 13 is a diagram schematically illustrating the difference between a voice-based natural language command and a multimodal (Vision-Language-Action, VLA) model in comparison to a conventional robot control method and a control system according to an embodiment of the present invention. Specific details for implementing the invention

[0033] Hereinafter, embodiments will be described in detail with reference to the attached drawings.

[0034] Embodiments of the present invention relate to the field of robot and physical AI technology, and more specifically, to a method for learning a Vision-Language-Action (VLA) model and a learning system (VLA model learning method and VLA model learning system) that predicts the next action by integrating visual information, language information, and state information of a robot, and furthermore, to a robot control method and a robot control system for generating and controlling the next action of a robot using the VLA model learned as above.

[0035] A VLA model learning system according to one embodiment may be implemented by at least one computer device. In this case, a computer program according to one embodiment may be installed and run on at least one computer device, and at least one computer device may perform a VLA model learning method according to one embodiment under the control of the run computer program. The above-described computer program may be stored on a non-transient computer-readable recording medium that is combined with at least one computer device to execute the VLA model learning method on the computer.

[0036] A robot control system according to one embodiment may be implemented by at least one computer device. In this case, a computer program according to one embodiment may be installed and run on at least one computer device, and at least one computer device may perform a robot control method according to one embodiment under the control of the run computer program. The above-described computer program may be stored on a non-transient computer-readable recording medium that is combined with at least one computer device to execute the robot control method on the computer.

[0037] FIG. 1 is a diagram illustrating an example of an overview of a Vision-Language-Action (VLA) based system according to one embodiment of the present invention. A VLA based system (100) according to this embodiment may be configured to include a robot work environment (110), which is a physical environment in which a robot performs a task; a robot (120) that performs the actual task; a VLA model learning system (130) for learning a VLA model; and a robot control system (140) that controls the robot using the learned VLA model. The robot work environment (110) is an area in a real or virtual space where the robot (120) performs tasks and collects data. It may include various observation devices such as a camera, a depth sensor, and a surrounding environment detection sensor, and may be configured to provide visual information and environmental information generated during the robot's work process in real-time or non-real-time. In addition, in the embodiments of the present invention, the work environment used in the learning phase and the work environment used in the actual robot control phase may differ from each other; for example, the learning environment for collecting robot motion data may be a data collection room equipped with fixed shooting equipment, whereas the robot control environment may be an actual industrial site or a service robot installation location.

[0038] The robot (120) can be implemented in various forms such as a multi-joint manipulator, a mobile robot, or a humanoid robot, and can be configured to generate state information such as its joint status, posture information, and motion history by including joint sensors, motor drivers, and end-effector position and posture sensors. The robot (120) can perform actions such as manipulating a real object, moving, and trajectory following by receiving and executing a driving command from the robot control system (140), and the state information generated during the operation process can be provided back to the robot control system (140) or the VLA model learning system (130).

[0039] The VLA model learning system (130) is a dedicated computing device for training a VLA model by integrating visual information, language information, and robot state information, and may include one or more processors and high-performance computing resources. The VLA model learning system (130) may operate to train a VLA model using a learning dataset composed of visual information collected in a robot work environment (110), state information and operation history of the robot (120), user or script-based language instructions, and multiple robot work episodes. The trained VLA model may be stored in the form of a file and transmitted to a robot control system (140), and in an embodiment of the present invention, the VLA model learning system (130) may be operated in a remote server or cloud environment that is physically separated from the location where the actual robot is installed.

[0040] The robot control system (140) can perform the function of generating the next action of the robot using a learned VLA model and providing it to the robot (120). The robot control system (140) may include a vision encoder that processes visual information acquired from the robot work environment (110), a language encoder that generates embeddings by understanding natural language instructions from a user, a state encoder that processes the joint status and motion history of the robot, a multimodal inference module that generates action tokens by integrating the embeddings, and a command converter that converts the action tokens into driving commands representing joint commands or end-effector trajectories of the robot. The robot control system (140) can be implemented in various forms such as a local controller, an edge device, or a server-based system, and can be connected to the robot (120) via a wired or wireless network or a dedicated communication link.

[0041] In addition, as illustrated in FIG. 1, the robot work environment (110), the robot (120), the VLA model learning system (130), and the robot control system (140) can be connected to each other via a network, and can also be configured to share a data storage or be linked with an external cloud. In one embodiment of the present invention, although the learning phase and the control phase may be performed in different hardware environments or different network environments, the two phases can be logically combined through the VLA model to enable consistent robot behavior control.

[0042] FIG. 2 is a block diagram illustrating an example of the internal configuration of a VLA model learning system according to an embodiment of the present invention, and FIG. 3 is a flowchart illustrating an example of a VLA model learning method according to an embodiment of the present invention.

[0043] The VLA model learning system (130) according to the present embodiment may include a visual information processing unit (210), a language information processing unit (220), a state information processing unit (230), a multimodal transformer model (240), a learning control unit (250), and a model management unit (260). The VLA model learning system (130) may be substantially implemented by at least one computer device, wherein the visual information processing unit (210), the language information processing unit (220), the state information processing unit (230), the multimodal transformer model (240), the learning control unit (250), and the model management unit (260) may be functional expressions for the operation of at least one processor included in at least one computer device.

[0044] The VLA model learning method according to the present embodiment may be performed by a VLA model learning system (130) implemented by at least one computer device. At this time, at least one processor included in the at least one computer device may be implemented to execute a control instruction according to the code of an operating system included in memory or the code of at least one computer program. Here, the at least one processor may operate according to a control instruction provided by the code stored in the at least one computer device to control the VLA model learning system (130) implemented by the at least one computer device so that the VLA model learning system (130) performs the steps (310 to 370) included in the method of FIG. 3.

[0045] In step (310), the VLA model learning system (130) can generate visual embeddings by inputting visual information acquired from a robot work environment into a vision encoder through a visual information processing unit (210). Here, the robot work environment may correspond to the robot work environment (110) described through FIG. 1. The visual information processing unit (210) can preprocess image data acquired from a camera or depth sensor placed in the robot work environment (110) and provide it to the vision encoder. The vision encoder can generate visual embeddings by extracting tokens in the form of image patches from the input visual information and then performing encoding based on a Vision Transformer. Additionally, the vision encoder can be fine-tuned to match the characteristics of the robot work environment based on a pre-trained model, thereby being configured to stably generate visual embeddings even under various visual conditions.

[0046] In step (320), the VLA model learning system (130) can generate language embeddings by inputting at least one of natural language instructions and text inputs into a language encoder through a language information processing unit (220). The language information processing unit (220) can receive natural language instructions provided by a user or a learning script, tokenize them, and provide them to the language encoder. The language encoder converts the input sentence into an embedding by utilizing the embedding space of a pre-trained Large Language Model (LLM), and the generated language embedding can clearly reflect the semantic intent of the task to be performed by the robot. Additionally, in the learning environment, data can be expanded to include instructions of various expression styles, so the language encoder can generate language embeddings with more robust generalization capabilities.

[0047] In step (330), the VLA model learning system (130) can generate state embeddings by inputting at least one of the robot's joint state, posture information, and motion history into a state encoder through the state information processing unit (230). The state information processing unit (230) can collect temporal state information, such as the robot's joint angles, joint velocities, end effector positions, and posture information, as well as the robot's past motion history, and normalize it to provide it to the state encoder. The state encoder generates state embeddings by performing a neural network-based state encoding process, and these state embeddings can reflect the robot's current range of motion and motion context. In this embodiment, these various state information are integrated to enable more accurate behavior prediction learning.

[0048] In step (340), the VLA model learning system (130) inputs visual embeddings, language embeddings, and state embeddings into the multimodal transformer model (240) through the learning control unit (250) to learn the correlations between the embeddings and generate action tokens. The multimodal transformer model (240) performs self-attention for each of the visual, language, and state embeddings, and can perform cross-attention to learn the semantic relationships between different embeddings. Additionally, position encoding can be additionally input to reflect the temporal sequence information of each embedding, thereby enabling the model to predict actions considering the flow of time. Through these structural features, the multimodal transformer model can generate action tokens that represent the robot's next action.

[0049] In step (350), the VLA model learning system (130) utilizes a learning dataset composed of multiple robot work episodes through the learning control unit (250), and may construct the learning dataset using at least one of (i) a uniform random sampling method that randomly selects the location of a target object, and (ii) a hierarchical sampling method that divides the workspace into layers and collects episodes by layer. The learning dataset composition unit (250) manages various work episodes collected in the robot work environment and can minimize compositional bias of the dataset by applying the uniform random sampling method or the hierarchical sampling method. The uniform random sampling method ensures data diversity by ensuring that the location of the target object or the arrangement of the environment is randomly selected, while the hierarchical sampling method can mitigate data imbalance by dividing the workspace based on region or difficulty and selecting an equal number of episodes for each layer. These strategies play an important role in improving the generalization performance of the model.

[0050] In step (360), the VLA model learning system (130) can obtain a learned VLA model that predicts the robot's next action by updating the parameters of the multimodal transformer model through the learning control unit (250) to minimize the difference between the action token and the robot's actual motion data. The learning control unit (250) can input training data organized in batches and calculate a loss based on the difference between the action token and the target action vector (ground-truth motion sequence). An optimizer (e.g., Adam, SGD, etc.) is used to minimize this loss, and the parameters of the multimodal transformer model can be progressively optimized through multiple learning epochs. This embodiment enables the acquisition of a VLA model capable of accurately predicting actual robot motion by sufficiently learning the correlation between the three modalities of vision, language, and state.

[0051] In this specification, the terms "optimize parameters" or "optimize" may refer to a series of numerical computational procedures for updating a set of parameters, including weights and biases, that constitute a model so as to minimize the loss function value between the predicted value output by the multimodal transformer model and the correct label. Such optimization may be performed, for example, using gradient-based optimization algorithms (e.g., Stochastic Gradient Descent (SGD), Adam, RMSProp, etc.) that calculate the gradient of the loss function and adjust parameters based on that gradient. In other words, optimization may refer to a series of computational procedures that improve the predictive performance of a model by iteratively adjusting parameter values ​​in a direction that reduces the loss value for an input batch. However, such optimization algorithms are not limited to specific types.

[0052] In step (370), the VLA model learning system (130) can store or provide the acquired VLA model through the model management unit (260). The model management unit (260) stores the learned VLA model in file form, performs version control functions, and can transfer the model to the robot control system (140) or an external server as needed. Additionally, intermediate learning results can be regularly stored and utilized for resume training or performance verification. The model transferred to the robot control system (140) can be directly used for generating robot behavior in the subsequent robot control stage.

[0053] FIG. 4 is a block diagram illustrating an example of the internal configuration of a robot control system according to an embodiment of the present invention, and FIG. 5 is a flowchart illustrating an example of a robot control method according to an embodiment of the present invention.

[0054] The robot control system (140) may include a vision encoder (410) configured to generate visual embeddings, a language encoder (420) configured to generate language embeddings, a state encoder (430) configured to generate state embeddings, a multimodal inference module (440) configured to generate behavior tokens, and a command converter (450) configured to convert behavior tokens into robot driving commands. The robot control system (140) may be implemented by one or more computer devices including one or more processors, and each component may be a functional module representing the functional operation of one or more processors.

[0055] A vision encoder (410) can be configured to receive visual information acquired from a robot work environment as input and generate visual embeddings. The vision encoder (410) can preprocess image data provided from a camera or other image sensor, extract tokens in the form of image patches based on this data, and then generate visual embeddings using a neural network model that includes a Vision Transformer (ViT) structure. This model can be fine-tuned to reflect the characteristics of the actual robot environment after undergoing pretraining. The generated visual embeddings can be provided to a multimodal inference module to utilize visual information such as the structure of the environment, the location of objects, and lighting conditions.

[0056] The language encoder (420) may be configured to receive at least one of the user's natural language instructions and text input as input to generate language embeddings. The language encoder (420) may tokenize speech recognition results and / or text directly entered by the user and convert them into an embedding space of a large language model. The language model may operate based on pre-trained language structures to understand various forms of natural language instructions. The generated language embeddings provide semantic context for tasks to be performed by the robot and enable the structural understanding of complex commands.

[0057] The state encoder (430) may be configured to generate a state embedding by receiving at least one of the robot's joint state, posture information, and motion history as input. The state encoder (430) may collect information such as the robot's joint angles, joint velocities, end-effector posture, and the robot's recent motion sequence, normalize it, and then perform neural network-based encoding. The generated state embedding can support more accurate behavior prediction by reflecting not only the robot's instantaneous physical state but also the context of the preceding action.

[0058] The multimodal inference module (440) can be configured to receive visual embeddings, language embeddings, and state embeddings as inputs and generate behavior tokens from a trained VLA model. The multimodal inference module (440) can perform self-attention operations for each of the visual, language, and state embeddings, and cross-attention operations to learn the relationship between modals. Position encoding may be applied during the processing, thereby generating behavior tokens that reflect the temporal order. The generated behavior tokens serve as the basis for generating low-level commands for robot control thereafter.

[0059] The command conversion unit (450) may be configured to convert an action token into a robot drive command representing a joint command or an end-effector trajectory of the robot. The command conversion unit (450) may include a decoding module that converts the action token into a continuous robot control command format. For example, the action token may be converted into a joint-space command or a task-space trajectory command in the end-effector reference coordinate system that can be directly transmitted to the robot's joint motors and provided to the robot device. This process may additionally include a post-processing step that takes into account the robot's dynamic stability and collision avoidance.

[0060] The robot control method according to the present embodiment may be performed by a robot control system (140) implemented by at least one computer device. At this time, at least one processor included in the at least one computer device may be implemented to execute a control instruction according to the code of an operating system included in memory or the code of at least one computer program. Here, the at least one processor may operate according to a control instruction provided by the code stored in the at least one computer device to control the robot control system (140) implemented by the at least one computer device so that the robot control system (140) performs steps (510 to 570) included in the method of FIG. 5.

[0061] In step (510), the robot control system (140) can generate visual embeddings by inputting visual information acquired from the robot work environment into a vision encoder. The visual information acquired in this step may be in various forms, such as RGB images, depth maps, and segmentation images, and the vision encoder can generate visual embeddings by converting them into tokens and performing converter-based encoding. These visual embeddings may include information essential for performing tasks, such as the structure of the robot's surrounding environment and the state of objects.

[0062] In step (520), the robot control system (140) can generate language embeddings by inputting at least one of the user's natural language instructions and text inputs into a language encoder. Natural language instructions can be converted into text by a speech recognition model, and the language encoder can process these text inputs using the structure of a large-scale language model. The generated language embeddings may include semantic information such as the purpose, requirements, and operation method of the task that the robot must perform. For example, the robot control system (140) can convert the user's voice input into text using a speech recognition model and input the converted text into a large-scale language model to generate control instructions for the robot to perform. This step is an extension of the language embedding generation process and can perform the function of automatically converting natural language voice input into commands that the robot can understand. The generated control instructions can be converted into an embedding form by the language encoder. The large-scale language model can be configured to perform inference using a system prompt to convert these natural language instructions into control commands in a format that the robot can execute. This involves utilizing system prompt templates to enable large-scale language models to generate instructions optimized for robot control, thereby allowing for the creation of structured output tailored to specific task objectives.

[0063] In step (530), the robot control system (140) can generate a state embedding by inputting at least one of the robot's joint state, posture information, and motion history into a state encoder. The state encoder can construct the state embedding by utilizing not only the robot's current posture and movement but also its previous motion history. Through this, the system can perform action inference considering what physical state the robot is currently in and what actions are possible.

[0064] In step (540), the robot control system (140) can generate action tokens by inputting visual embeddings, language embeddings, and state embeddings into a trained VLA model. Since the trained VLA model has already learned semantic correlations between multimodal inputs, it can generate action tokens representing the next action based on the input embeddings. At this time, position encoding may be included, making it possible to perform action inference that reflects the temporal order.

[0065] In step (550), the robot control system (140) can convert the action token into a robot drive command representing a joint command or end-effector trajectory of the robot and provide it to the robot device to control the robot to perform the next action. This conversion may include action decoding, robot motion planning, dynamic correction, etc., and the generated motion command can be transmitted to the robot's control board or motor driver to execute the actual action.

[0066] In step (560), the robot control system (140) can periodically collect sensor feedback of the robot, generate a new state embedding reflecting the sensor feedback, and then repeatedly execute a learned VLA model using the new state embedding to update the robot's behavior in real time. The sensor feedback may include joint errors, external force sensor information, vision-based tracking information, etc., and can respond to real-time control by continuously updating the robot's behavior plan by reflecting this.

[0067] FIG. 6 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. For example, the VLA model learning system (130) and the robot control system (140) described above may each be implemented by at least one computer device. In this case, each of the at least one computer device may correspond to the computer device (600) of FIG. 6. As shown in FIG. 6, the computer device (600) may include memory (610), a processor (620), a communication interface (630), and an input / output interface (640). The memory (610) is a computer-readable recording medium and may include a non-perishable mass storage device such as RAM (random access memory), ROM (read only memory), and a disk drive. Here, non-perishable mass storage devices such as ROM and disk drives may be included in the computer device (600) as separate permanent storage devices distinct from memory (610). Additionally, an operating system and at least one program code may be stored in memory (610). These software components may be loaded into memory (610) from a computer-readable recording medium separate from memory (610). This separate computer-readable recording medium may include computer-readable recording media such as floppy drives, disks, tapes, DVD / CD-ROM drives, and memory cards. In another embodiment, software components may be loaded into memory (610) via a communication interface (630) rather than a computer-readable recording medium.For example, software components can be loaded into the memory (610) of a computer device (600) based on a computer program installed by files received through a network (Network, 660).

[0068] The processor (620) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (620) via memory (610) or a communication interface (630). For example, the processor (620) may be configured to execute instructions received according to program code stored in a recording device such as memory (610).

[0069] The communication interface (630) may provide a function for the computer device (600) to communicate with other devices through the network (660). For example, requests, commands, data, files, etc. generated by the processor (620) of the computer device (600) according to program code stored in a recording device such as memory (610) may be transmitted to other devices through the network (660) under the control of the communication interface (630). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (600) through the communication interface (630) of the computer device (600) via the network (660). Signals, commands, data, etc. received through the communication interface (630) may be transmitted to the processor (620) or memory (610), and files, etc. may be stored in a storage medium (the permanent storage device described above) that the computer device (600) may further include.

[0070] The input / output interface (640) may be a means for interfacing with an input / output device (I / O device, 650). For example, the input device may include a device such as a microphone, keyboard, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (640) may be a means for interfacing with a device in which the functions for input and output are integrated into one, such as a touchscreen. The input / output device (650) may be composed of a computer device (600) and a single device.

[0071] Additionally, in other embodiments, the computer device (600) may include fewer or more components than those of FIG. 6. However, it is not necessary to clearly illustrate most of the prior art components. For example, the computer device (600) may be implemented to include at least some of the input / output devices (650) described above, or may include other components such as a transceiver, a database, etc.

[0072] FIG. 7 is a schematic diagram illustrating the input processing and action generation processes of a Vision-Language-Action (VLA) model according to an embodiment of the present invention. As illustrated in FIG. 7, the information flow for robot control can begin with three inputs: visual observation, language instruction, and robot state. First, visual observation information acquired from the robot work environment is input into a vision encoder to extract feature tokens at the image patch level and convert them into vision embeddings. The user's language instruction can be converted into text embeddings through a process of tokenization and mapping to an embedding space by a text encoder. Additionally, robot state information, consisting of the robot's joint angles, joint velocities, end-effector posture, etc., is input into a state encoder and converted into state embeddings representing the corresponding state.

[0073] Vision embeddings, text embeddings, and state embeddings are input into a multimodal transformer, which can integrally learn the associations between visual, linguistic, and state information by performing self-attention within each embedding and mutual attention among the embeddings. Based on this integrated representation, the multimodal transformer can output behavior tokens to a behavior token generator. The behavior token generator can generate behavior tokens representing the next action the robot must perform by reflecting the temporal order and semantic relationships of the multimodal representation.

[0074] Finally, the generated action token is passed to an action decoder and can be converted into a robot drive command representing the actual joint commands of the robot or the continuous trajectory of the end effector. The robot drive command generated by the action decoder is passed to a robot controller and can be executed as the physical movement of the robot.

[0075] As such, FIG. 7 clearly illustrates the overall structural flow of the VLA-based robot control system of the present embodiment, which integrates visual information, language instructions, and robot state information in a multimodal manner to generate robot behavior.

[0076] FIG. 8 is a diagram illustrating an example of the overall processing flow in which a continuous control command of a robot is generated from a voice-based natural language input through a multimodal visual-language-action (VLA) model according to an embodiment of the present invention. As shown in FIG. 8, the voice-based robot control system according to the present embodiment may be configured to perform the entire pipeline of generating a continuous robot action from a voice input in stages.

[0077] First, the user's voice input can be acquired through a microphone or an external voice acquisition device and provided to a speech recognition model (STT Model) to be converted into input text. At this stage, the speech recognition model can generate input for the subsequent language understanding stage by accurately converting the spoken natural language command into text.

[0078] The converted text can be passed to a large-scale language model (LLM with Instructions), which can reconstruct the input sentence into structured language instructions suitable for robot execution based on system prompt-based rules. These instructions can be converted to clearly include the actions the robot must perform, target objects, locations, constraints, and more.

[0079] Next, within the VLA model, three modalities—visual information, language information, and state information—can be processed in parallel.

[0080] (1) Vision information can be input from sensors such as cameras and can be converted into vision embeddings through a vision encoder.

[0081] (2) Language instructions generated by the LLM can be converted into language embeddings by a language encoder.

[0082] (3) The robot's posture, joint angles, end effector state, etc. can be converted into a low-dimensional representation called a state embedding by a state encoder.

[0083] The three types of embeddings generated in this way are input into a Multimodal Transformer, which can integrally learn the interrelationships between visual, language, and state information through self-attention and cross-attention-based operations. As a result, a Continuous Action Output representing the actions the robot must perform can be generated.

[0084] The generated continuous control commands are transmitted to the robot device, which can perform physical actions such as actual joint control or end-effector movement. Additionally, the robot's state is fed back into the system and reflected in the next action generation process, thereby enabling real-time responsive control.

[0085] As such, FIG. 8 clearly illustrates the entire processing path of the present embodiment leading from voice input to text conversion, language instruction generation, visual, language, and state-based multimodal integrated inference, and robot behavior output, and clearly explains the technical features of the system of the present embodiment that convert complex natural language commands into actual robot behaviors and the interactions between the components.

[0086] FIG. 9 is a diagram illustrating screens showing an actual implementation example of a robot control system according to an embodiment of the present invention. In the embodiment of FIG. 9, the process of a robot picking up a designated object and delivering it to the user based on the user's voice input is described sequentially. In this embodiment, the robot control system includes a voice recognition model, a language encoder, a state encoder, a vision encoder, a trained VLA model, and an action decoder, and can receive and interpret natural language commands from the user in real time to generate and execute the robot's next action.

[0087] The first screen of FIG. 9 shows a scene where a user utters the natural language voice input, "Pick up the banana and hand it over." The voice input is converted into text by a speech recognition model, and then converted into language embeddings through a process of generating control instructions via a language encoder and a large-scale language model (LLM). At the same time, an image captured in the robot working environment is input into a vision encoder and converted into visual embeddings, and the robot's current joint state and posture information can be converted into state embeddings by a state encoder.

[0088] The second screen of FIG. 9 shows a scene in which a robot performs the action of picking up a banana, which is the work target, based on an action token generated by a robot control system. The trained VLA model generates an action token that the robot must perform next based on input visual embeddings, language embeddings, and state embeddings, and the action decoder can convert the action token into a series of robot drive commands representing the robot's joint control commands or end-effector trajectories. The robot can be controlled to stably grasp the banana at an accurate position by executing the converted robot drive commands.

[0089] The third screen of FIG. 9 shows a scene in which a robot moves a banana it has picked up toward a user and finally delivers it to the user. During this process, the robot control system updates the state embeddings based on sensor feedback from the robot that is collected periodically, and corrects the robot's movement in real time by repeatedly executing a learned VLA model using the updated state embeddings. Accordingly, the robot can stably deliver the banana to the user even if there are environmental changes, such as changes in the position of an object or slight movements of the user's position.

[0090] As such, the embodiment of FIG. 9 clearly illustrates the entire process of a robot control system interpreting voice-based natural language commands in real time and controlling an actual robot using a VLA model that integrates visual, language, and robot state information, and demonstrates that consistent and accurate robot movements can be performed even for complex work commands.

[0091] The VLA model according to the present embodiment can perform color and shape-based semantic object detection in complex environments where multiple objects exist, rather than at a simple pick-and-place level. For example, tasks such as "putting a sandwich in a gray box" or "putting a can in a brown box" are high-difficulty tasks that simultaneously require color-based box identification and object type differentiation. The pre-trained SmolVLA model can perform operations based on semantic differentiation even in such scenarios, which demonstrates the advantages of the multimodal learning structure provided by the present invention.

[0092] FIG. 10 illustrates an example of the overall processing flow in which a voice-based natural language command is converted into an executable control command for a robot in a robot control system according to an embodiment of the present invention. As shown in FIG. 10, when a user utters a command in natural language, a speech recognition model receives the corresponding voice input and can output the command in text form. Subsequently, the command in text form is transmitted to a large language model to convert the task to be performed by the robot into a more structured command form, and the converted command can be transmitted as a robot control command. The robot control command is then input into a visual-language-behavioral model (VLA model), whereby an action is generated by reflecting the visual information of the actual work environment and the robot state, and finally, the robot device can perform the task based on the corresponding action command. FIG. 10 provides an overview of the entire pipeline leading from voice to text to language model to control command to action model to task execution, demonstrating that the robot control process according to the present embodiment is organically connected from natural language input to the execution of physical actions.

[0093] FIG. 11 is a diagram illustrating an example of the process in which a text-based command is converted into a robot control command in a robot control system according to an embodiment of the present invention. As shown in FIG. 11, when a user inputs a command in text form, the system transmits the text command over a network via an API call, and an external large-scale language model can reconstruct the command into a shortened control sentence form that the robot can execute. The large-scale language model analyzes the input command based on the system prompt to generate a formalized control sentence including a clear action, object, and location / target, and the generated sentence can be transmitted as a robot control command after undergoing an output post-processing step. FIG. 11 clearly explains the feature of increasing the accuracy of robot action decisions by converting natural language commands into short, imperative control sentences that the robot can execute. Additionally, the bottom of the figure includes an example of an API prompt used by the large-scale language model, illustrating an example of the rules for converting natural language sentences into robot control commands.

[0094] FIG. 12 illustrates an example of the internal information processing flow of a visual-language-action model that generates continuous actions based on visual, language, and state information in a robot control system according to an embodiment of the present invention. As shown in FIG. 12, natural language commands can be converted into text through a speech recognition model and a large language model, refined into robot control commands, and provided as input to a VLA model. At the same time, visual observations of the robot work environment are converted into vision embeddings through a vision encoder, and state information such as the robot's joint angles, velocity, and posture can be converted into state embeddings through a state encoder. A language encoder can convert text commands into text embeddings. The generated vision embeddings, text embeddings, and state embeddings are integrally input into a multimodal transformer, and the multimodal transformer can learn the interrelationships between each piece of information through self-attention and cross-attention operations. Subsequently, the action token generation module outputs an action token representing the next action, and the action decoder converts the token into a robot control command in the form of a continuous action vector that the robot can execute, and transmits it to the robot device. Figure 12 specifically illustrates an example of a technology flow that generates continuous robot actions in real time by integrating multimodal information of vision, language, and state.

[0095] Meanwhile, one of the primary motivations for developing the embodiments of the present invention is to resolve the "embodiment gap problem" that occurs between different robot platforms. The embodiment gap refers to a phenomenon where motion failure occurs when a policy learned based on the joint structure, degrees of freedom, and dynamic characteristics of a specific robot is applied to a different type of robot. This manifests as a situation where a pre-trained SmolVLA policy successfully performs object grasping tasks on SO-100 series robots, but fails to execute the final grasp motion even if the object's position is recognized on robots with different kinematic structures, such as the WidowX-250.

[0096] To solve the Embodiment Gap problem, the VLA model learning system and robot control system according to embodiments of the present invention may use state embeddings including the actual joint state and motion history of the robot, and provide a structure that integrally learns the correlation between visual, linguistic, and state information based on a multimodal transformer. Through this, generalization performance can be improved so that robots can perform instructed tasks in a consistent manner even if they have different body structures.

[0097] Although pre-trained VLA models are trained based on various public datasets, domain gaps may occur in actual robot work environments due to differences in camera position, background, and lighting conditions. In particular, smaller models are subject to the constraint that "the camera position during training and the camera position during inference must be the same," and using the pre-trained model as is can lead to significant performance degradation. To address these domain gaps, embodiments of the present invention may include a process of fine-tuning the model based on actual work episodes performed by the robot itself. Through this, the pre-trained model adapts to new work environments, and experimental results have confirmed that a performance improvement of approximately 25% or more is possible.

[0098] FIG. 13 is a diagram schematically illustrating the difference between a voice-based natural language command and a multimodal (Vision-Language-Action, VLA) model in comparison to a conventional robot control method and a control system according to an embodiment of the present invention.

[0099] The conventional method illustrated at the top of Fig. 13 represents an example in which a direct control method using a teaching pendant or a code-based control method by a professional programmer is used to control a robot. In this conventional method, since the user must directly specify the actions to be performed by the robot using coordinates or program the robot motion sequence in the form of code, a high level of expertise is required, and there are problems such as long preparation time and poor adaptability to changes in the environment.

[0100] The proposed control system of the present invention, illustrated at the bottom of FIG. 13, represents a series of flows that begin with a user's voice input and sequentially pass through speech recognition (STT: Speech-To-Text), a large-scale language model (LLM), and a visual-language-action (VLA) model to ultimately automatically generate action commands for a robot. The user simply needs to utter a simple natural language command, the speech recognition model converts it into text, and the LLM reinterprets the text into structured instructions that the robot can perform. Subsequently, the VLA model integrates visual information, language information, and state information to sequentially generate the robot's next action.

[0101] As such, FIG. 13 visually demonstrates the technical differentiation that, unlike conventional manual and expert-centered control methods, the control structure of the present invention enables even non-experts to control a robot using only intuitive and natural language-based commands.

[0102] As described above, according to the embodiments of the present invention, a learning method for a Visual-Language-Action (VLA) model that integrally processes visual information, language information, and robot state information, as well as a robot control method and system, are provided. This enables the intuitive interpretation of a user's natural language commands and the stable generation of the robot's subsequent corresponding actions, even in complex work environments. In particular, a multimodal transformer structure based on pre-trained vision, language, and state encoders efficiently learns the correlations between each piece of information, thereby ensuring high predictive performance even in a lightweight structure. Furthermore, a dataset construction method applying uniform random sampling or hierarchical sampling can improve model generalization by reflecting various task distributions in a balanced manner. Additionally, the trained VLA model can continuously update the robot's actions by reflecting real-time sensor feedback, enabling stable operation even in dynamic environments. Moreover, by providing a natural language-based robot control interface through integration with a speech recognition model and a large-scale language model, the user experience can be significantly enhanced. Ultimately, according to the embodiments of the present invention, high-performance multimodal action policies can be implemented even on low-cost robot platforms, making it highly useful for building practical and scalable physical AI systems.

[0103] The system or device described above may be implemented as a hardware component, or a combination of a hardware component and a software component. For example, the device and component described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0104] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0105] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components; it is not limited to a medium directly connected to a computer system, but may also exist distributed over a network. Examples of computer-readable storage media include ROM, PROM, EPROM, EEPROM, flash memory (e.g., NAND / NOR), SSD, HDD, magnetic tape, optical recording media (CD-ROM, DVD, BD), magneto-optical media, memory cards, USB memory, etc. Such storage media refer to non-transient media and do not include transmission signals themselves or purely volatile memory (e.g., RAM). In addition, other examples of media include recording or storage media managed by app stores that distribute applications, or sites and servers that supply or distribute various other software. Examples of program instructions include not only machine code, such as that generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc.

[0106] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0107] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

Claim 1 A method for training a VLA model in a VLA model learning system that trains a Vision-Language-Action (VLA) model that predicts the next action of a robot by integrating visual information, language information, and state information of the robot, wherein the VLA model learning system is implemented by at least one computer device including at least one processor, and the VLA model learning method comprises: a step of generating a visual embedding by inputting visual information acquired from a robot work environment into a vision encoder by the at least one processor; a step of generating a language embedding by inputting at least one of natural language instruction and text input into a language encoder by the at least one processor; and a step of generating a state embedding by inputting at least one of the robot's joint state, pose information, and motion history into a state encoder by the at least one processor. A step of inputting the visual embeddings, language embeddings, and state embeddings into a multimodal transformer model by the at least one processor, thereby learning the correlations between the embeddings and generating action tokens by performing self-attention for each embedding and cross-attention between the embeddings; a step of constructing the training dataset by the at least one processor using a training dataset composed of multiple robot task episodes, wherein the training dataset is constructed using at least one sampling method among (i) a uniform random sampling method that randomly selects the location of a target object, and (ii) a hierarchical sampling method that divides the workspace into layers and collects episodes by layer, wherein the at least one sampling method is performed to reduce bias in the data distribution;A method for learning a VLA model comprising the step of obtaining a learned VLA model that predicts the next action of a robot by updating the parameters of the multimodal transformer model based on the difference between the action token and the actual motion data of the robot by the at least one processor so as to reduce the difference. Claim 2 A learning method for a VLA model according to claim 1, characterized in that the hierarchical sampling method is performed to reduce bias in the data distribution by dividing the workspace into multiple layers and collecting an equal number of episodes from each layer. Claim 3 delete Claim 4 A method for learning a VLA model according to claim 1, wherein the parameter update of the multimodal transformer model is performed based on a loss function set to minimize the difference between the action token and a target action vector representing the actual continuous motion of the robot. Claim 5 A method for learning a VLA model according to claim 1, wherein the vision encoder, language encoder, and state encoder are each fine-tuned based on a pre-trained model, and the fine-tuning is performed using visual information captured in a robot work environment and collected robot motion data. Claim 6 A method for training a VLA model according to claim 1, wherein the step of generating the visual embedding is characterized by extracting tokens in the form of image patches from input visual information and converting the tokens into embedding vectors using vision converter-based encoding. Claim 7 A method for training a VLA model according to claim 1, wherein the step of generating the language embedding is characterized by generating the language embedding using a language encoding that tokenizes natural language instructions and maps the tokens to the embedding space of a pre-trained language model. Claim 8 A learning method for a VLA model according to claim 1, wherein the step of generating the state embedding is characterized by receiving at least one of the robot's joint angle, joint velocity, and end effector pose as input, and generating the state embedding by projecting the value received as input into an embedding space using a state encoding neural network. Claim 9 A method for learning a VLA model according to claim 1, wherein the step of generating the action token is characterized by generating the action token by reflecting the sequential meaning of the embeddings by further inputting a position encoding representing the temporal order of each embedding into a multimodal transformer model in addition to the visual embedding, language embedding, and state embedding. Claim 10 A method for training a VLA model according to claim 1, wherein the parameter update of the multimodal transformer model is performed using a training epoch process in which a training dataset is input in batch units and the parameters are updated based on the loss value for said batch. Claim 11 A robot control method for generating the next action of a robot using a learned Vision-Language-Action (VLA) model, wherein the robot control method is executed by a computer device comprising at least one processor, and wherein the at least one processor inputs visual information acquired from a robot work environment into a vision encoder to generate a visual embedding; wherein the at least one processor inputs at least one of a user's natural language instruction and text input into a language encoder to generate a language embedding; wherein the at least one processor inputs at least one of the robot's joint state, posture information, and motion history into a state encoder to generate a state embedding; wherein the at least one processor inputs the visual embedding, language embedding, and state embedding into the learned VLA model to generate an action token that reflects the correlation between the embeddings based on self-attention for each embedding and cross-attention between the embeddings. A robot control method characterized by including the step of converting the action token into a robot driving command representing a joint command or an end-effector trajectory of the robot by the at least one processor, and providing the robot driving command to a robot device to control the robot to perform the next action. Claim 12 A robot control method according to claim 11, wherein the step of generating the language embedding is characterized by converting a user's voice input into text using a speech recognition model, inputting the text into a large-scale language model to generate control instructions to be performed by the robot, and inputting the control instructions into the language encoder to generate the language embedding. Claim 13 A robot control method according to claim 12, wherein the large-scale language model performs inference using a system prompt to convert natural language instructions into control commands in a format executable by the robot. Claim 14 A robot control method according to claim 11, wherein the robot control method further comprises the step of periodically collecting sensor feedback of the robot by the at least one processor, generating a new state embedding reflecting the sensor feedback, and then repeatedly executing the learned VLA model using the new state embedding to update the robot's behavior in real time. Claim 15 A robot control system for generating the next action of a robot using a learned Vision-Language-Action (VLA) model, comprising: a vision encoder configured to generate a visual embedding by receiving visual information acquired from a robot work environment as input; a language encoder configured to generate a language embedding by receiving at least one of a user's natural language instruction and text input as input; a state encoder configured to generate a state embedding by receiving at least one of the robot's joint state, pose information, and motion history as input; and a multimodal inference module configured to generate an action token reflecting the correlation between the embeddings by receiving the visual embedding, language embedding, and state embedding as inputs and performing self-attention for each embedding and cross-attention between the embeddings. A robot control system comprising a command converter configured to convert the action token into a robot driving command representing a joint command or an end-effector trajectory of the robot, wherein the robot is configured to perform the next action by receiving and executing the robot driving command generated by the command converter.

Citation Information

Patent Citations

  • Multi-modal feedback integration type knowledge-search enhanced robot control system

    JP2025108598A

  • Training and / or utilizing machine learning model(s) for use in natural language-based robot control.

    KR1020230008171A

  • Method for training artificial intelligence model and apparatus for the same

    KR1020250043970A

  • Method, apparatus, and recording medium for generating robot model dataset using artificial intelligence

    KR102841574B1