Robot locomotion mode determination method and system

JP2026527625APending Publication Date: 2026-08-14NAVER CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0009】 本開示の一部の実施形態によれば、ゼロショットベースの視覚言語モデルを用いることで、新たな動的オブジェクトが存在する場合でも、ロボットが学習することなく状況を理解することができる。また、ロボットは、現在の状況を認識して柔軟に走行モードを切り替えることができる。これにより、ロボットの走行効率を向上させることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026527625000001_ABST
    Figure 2026527625000001_ABST
Patent Text Reader

Abstract

This disclosure provides a robot travel mode determination method, which is performed by at least one processor of the robot. The method includes the steps of: sending an image captured by the robot's camera and a query about the image to a server; receiving a response to the query from the server based on a visual language model; and determining the robot's travel mode based on the response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , , , , , , , ,

[0005] , , ,

[0003]

[0001] The present disclosure relates to a method and system for determining a robot running mode. Specifically, the present disclosure relates to a method and system for determining a robot running mode based on responses to queries generated by using a zero-shot based visual language model for an image captured by a robot and a query regarding the image, which are transmitted to a server.

Background Art

[0002] An autonomous mobile robot may refer to a robot that can autonomously move by recognizing the surrounding environment using a radar, LIDAR (light detection and ranging), GPS, a camera, or the like. In recent years, autonomous mobile robot services for providing services to people have been developed and utilized in various forms using autonomous mobile robot technology.

[0003] On the other hand, an ideal running method of a robot may vary depending on the characteristics of the environment. For example, when there are dynamic objects such as people, it is necessary to move while avoiding the dynamic objects, and in an environment where only robots are present, it is necessary to accurately move along a path set so as not to collide with each other. The conventional running method uses a method of moving in a specific running method in a specific area. However, such a conventional method cannot flexibly switch modes according to the situation and has limitations in running efficiency.

Summary of the Invention

Problems to be Solved by the Invention

[0006] According to one embodiment of the present disclosure, a robot driving mode determination method may include the steps of sending an image captured by the robot's camera and a query regarding the image to a server, receiving a response to the query from the server based on a visual language model, and determining the robot's driving mode based on the response.

[0007] A computer-readable non-temporary recording medium may be provided that stores instructions for executing a method according to one embodiment of the present disclosure on a computer.

[0008] According to one embodiment of the present disclosure, a robot is provided. The robot includes a communication module, memory, a display, and at least one processor connected to the memory and configured to run at least one computer-readable program contained in the memory, the at least one program may include instructions for sending images taken by the robot's camera and queries about the images to a server, receiving responses to the queries from the server based on a visual language model, and determining the robot's driving mode based on the responses. [Effects of the Invention]

[0009] According to some embodiments of this disclosure, by using a zero-shot-based visual language model, the robot can understand the situation without learning, even when new dynamic objects are present. Furthermore, the robot can recognize the current situation and flexibly switch between different modes of movement. This improves the robot's movement efficiency.

[0010] According to some embodiments of this disclosure, a robot can understand its surroundings using a zero-shot-based visual language model applicable to real-world environments, without learning new dynamic objects. This allows the robot to improve its mobility and reduce the risk of collisions by determining and switching to an appropriate mode of travel depending on the situation.

[0011] According to some embodiments of this disclosure, a zero-shot-based visual language model can be used to appropriately switch the robot's driving mode even when it is moving through a new environment.

[0012] According to some embodiments of this disclosure, real-time driving mode switching is possible by communicating with a server using a zero-shot-based visual language model.

[0013] The effects of this disclosure are not limited to those mentioned above, and any other effects not mentioned above would be clearly understood by a person with ordinary skill in the art to which this disclosure pertains ("ordinary art") from the wording of the claims. [Brief explanation of the drawing]

[0014] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, where similar reference numbers refer to similar elements, but are not limited thereto. [Figure 1] An example of a robot driving mode determination method according to one embodiment of this disclosure is shown. [Figure 2] This is a schematic diagram showing a configuration in which multiple robots and information processing systems are connected so that they can communicate with each other via a network. [Figure 3] This is a block diagram showing the internal configuration of a robot and an information processing system according to one embodiment of the present disclosure. [Figure 4] This figure shows an example of a driving mode determination factor according to one embodiment of the present disclosure. [Figure 5]A diagram showing an example in which a running mode of a robot is determined according to an embodiment of the present disclosure. [Figure 6] A diagram showing an example of a response table according to an embodiment of the present disclosure. [Figure 7] A diagram showing an example in which a running mode of a robot is determined according to an embodiment of the present disclosure. [Figure 8] A diagram showing an example in which a running mode of a robot is determined according to an embodiment of the present disclosure. [Figure 9] A diagram showing an example in which a running mode of a robot is determined according to an embodiment of the present disclosure. [Figure 10] A flowchart showing an example of a method according to an embodiment of the present disclosure.

Best Mode for Carrying Out the Invention

[0015] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, when the gist of the present disclosure may be unnecessarily obscured, specific descriptions of widely known functions and configurations are omitted.

[0016] In the accompanying drawings, the same or corresponding components are denoted by the same reference numerals. Also, in the following description of the embodiments, redundant descriptions of the same or corresponding components may be omitted. However, even if the description of a component is omitted, it is not intended that the component is not included in any of the embodiments.

[0017] The advantages and features of the disclosed embodiments, and the methods for achieving them, will become clear by referring to the embodiments described below together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and can be realized in various different forms. Merely, the present embodiments are provided to make the present disclosure complete and to enable those of ordinary skill in the art to fully understand the scope of the invention.

[0018] This specification provides a brief explanation of the terminology used herein and a detailed description of the disclosed embodiments. The terminology used herein has been selected to the greatest extent possible from currently widely used and common terms, taking into account the function of this disclosure; however, this may change due to the intent of the articulates in the relevant field, case law, the emergence of new technologies, etc. In some cases, the applicant has arbitrarily selected terms, in which case their meaning will be described in detail in the corresponding description of the invention. Therefore, the terms used in this disclosure should not be merely nouns, but should be defined based on the meaning of the term and the overall context of this disclosure.

[0019] In this specification, singular expressions include plural expressions unless the context clearly indicates that they are singular. Conversely, plural expressions include singular expressions unless the context clearly indicates that they are plural. Throughout the specification, where a part is described as containing a component, this does not exclude other components, unless otherwise stated.

[0020] Also, the terms "module" or "unit" used in the specification mean software or hardware components, and the "module" or "unit" performs a certain role. However, the "module" or "unit" does not mean being limited to software or hardware. The "module" or "unit" may be configured to be in an addressable storage medium or may be configured to cause one or more processors to reproduce. Thus, as an example, the "module" or "unit" may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The functions provided by the components and the "module" or "unit" can be combined into a smaller number of components and the "module" or "unit", or can be further separated into additional components and the "module" or "unit".

[0021] According to one embodiment of the present disclosure, a “module” or “part” can be realized by a processor and memory. “Processor” should be broadly interpreted to include general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, and the like. In some environments, “processor” may also refer to application-specific semiconductors (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), and the like. “Processor” may also refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other combination of such configurations. “Memory” should also be broadly interpreted to include any electronic component capable of storing electronic information. The term "memory" can also refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, and registers. Memory is said to be in electronic communication with the processor if the processor can read information from and / or record information into it. Memory integrated into a processor is in electronic communication with the processor.

[0022] In this disclosure, “System” may include, but is not limited to, at least one of a server device and a cloud device. For example, a system may consist of one or more server devices. Another example is that a system may consist of one or more cloud devices. Yet another example is that a system may consist of both server devices and cloud devices.

[0023] In this disclosure, "server," "system," and "information processing system" refer to devices capable of communicating with multiple robots, and may be used interchangeably with each other.

[0024] In this disclosure, “Display” may refer to any display device relating to a computing device, for example, any display device that is controlled by or provided by a computing device and is capable of displaying any information / data.

[0025] In this disclosure, “each of the A” or “each of the A” may refer to each of all components included in the A, or to each of some of the components included in the A.

[0026] In this disclosure, “machine learning model” may include any model used to infer an answer to a given input. According to one embodiment, a machine learning model may include an artificial neural network model comprising an input layer, a plurality of hidden layers, and an output layer, where each layer may include a plurality of nodes. In this disclosure, each of the plurality of machine learning models is described as a separate machine learning model, but is not limited to this, and some or all of the plurality of machine learning models can be implemented as a single machine learning model. Also, a single machine learning model may include multiple machine learning models. In this disclosure, the terms machine learning model and artificial neural network model may be used interchangeably to refer to the same or similar models.

[0027] In this disclosure, “visual language model” may refer to a machine learning model or artificial neural network model configured to understand and process visual data such as images. The visual language model may be pre-trained to describe visual information extracted from images or videos in natural language, or to generate visual content using natural language descriptions.

[0028] In this disclosure, “zero-shot-based visual-language models” may refer to visual-language models that learn from very small amounts of labeled or unlabeled data. Zero-shot-based visual-language models can learn about new domains and tasks even with unlabeled data by extracting pre-trained features for visual-language tasks using a pre-trained model and applying these features to new visual-language tasks.

[0029] Figure 1 shows an example of a robot driving mode determination method according to one embodiment of the present disclosure. In one embodiment, the robot 110 can autonomously travel to a destination 120. Specifically, the robot 110 can determine a travel path to the destination 120 based on absolute position information and autonomously travel along the travel path.

[0030] In one embodiment, the robot 110 can send images of the surrounding environment captured by a camera while moving, and queries related to the images, to a server. Here, the queries include queries related to factors determining the movement mode. For example, factors determining the movement mode may include the presence or absence of pedestrians, the presence or absence of other robots, the presence or absence of landmarks, whether the width of the passage is narrower than a predetermined value, and whether the social distance between the robot and a person is smaller than a predetermined value.

[0031] In one embodiment, the server can use a zero-shot-based visual language model 130 to generate a response to a query based on the image and query received from the robot 110. For example, a response may be generated indicating that the robot should travel at a lower speed than the reference speed because there are pedestrians 140 and other robots 150 in the image received from the robot 110. Examples of such responses will be described in detail later with reference to Figures 5 to 8.

[0032] In one embodiment, the response may be generated based on the behavior of a specific object within the image. For example, if a person in the image is sitting, that person is not a dynamic object and can be excluded from consideration. Alternatively, the response may be generated based on the characteristics of a specific object within the image. For example, the visual language model 130 may recognize that a person in the image is elderly or an infant and generate a response indicating that the vehicle should be driven in a pre-set driving mode (e.g., slow mode).

[0033] In one embodiment, the robot 110 can determine a driving mode based on a response received from a server. In this case, the robot 110 can determine one of a plurality of pre-set driving modes. Here, each of the plurality of driving modes may differ in at least one of the following: the robot's driving speed or the robot's autonomous driving tolerance level. For example, the plurality of pre-set driving modes may include, but are not limited to, an autonomous driving mode that drives at a speed lower than a reference speed, a strict driving mode that drives along a set path, a follow mode that moves along a wall, and a fast mode that drives at a speed higher than a reference speed, and may include various modes that combine driving speed and autonomous driving tolerance levels, such as a fast strict driving mode. Furthermore, the robot 110 can also determine a driving mode if it receives the same response from the server a predetermined number of times (e.g., 10 times).

[0034] Through this configuration, using a zero-shot-based visual language model, the robot can understand the situation without learning, even when new dynamic objects are present. Furthermore, the robot can recognize its current situation and flexibly switch between different travel modes. This improves the robot's travel efficiency. Additionally, by determining the travel mode only after receiving the same response a predetermined number of times, the robot can more accurately determine the travel mode.

[0035] Figure 2 is a schematic diagram showing a configuration in which multiple robots 210_1, 210_2, 210_3 and an information processing system 230 are connected so as to be able to communicate via a network 220. The information processing system 230 corresponds to a server according to one embodiment and may be configured to control the movement and / or operation of the multiple robots 210_1, 210_2, 210_3 via the network 220. According to one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more distributed computing devices and / or distributed databases based on a cloud computing service, that can store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to the control of the movement and / or operation of the multiple robots 210_1, 210_2, 210_3. The information processing system 230 may be located inside or outside the building in which the robots 210 are located.

[0036] Multiple robots 210_1, 210_2, and 210_3 can travel inside a building and communicate with an information processing system 230 via a network 220. The network 220 may be configured to enable communication between the multiple robots 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may consist of wired networks such as Ethernet (registered trademark), Power Line Communication, telephone line communication equipment, and RS-serial communication, wireless networks such as mobile communication networks, WLAN (Wireless LAN), Wi-Fi, Bluetooth (registered trademark), and ZigBee (registered trademark), or a combination thereof. The communication method is not limited and can include not only communication methods that utilize communication networks that network 220 may include (for example, mobile communication networks, wired internet, wireless internet, broadcasting networks, satellite networks, etc.), but also short-range wireless communication between multiple robots 210_1, 210_2, and 210_3. For example, network 220 may include one or more of the following networks: PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Network 220 may also include, but is not limited to, one or more network topologies, including bus networks, star networks, ring networks, mesh networks, starbus networks, tree-type or hierarchical networks, etc.

[0037] In Figure 2, the multiple robots 210_1, 210_2, and 210_3 may be any robots capable of wireless communication and autonomous navigation. Also, although Figure 2 shows three robots 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, it is not limited to this configuration, and a different number of robots 210_1, 210_2, and 210_3 may be configured to communicate with the information processing system 230 via the network 220.

[0038] According to one embodiment, the information processing system 230 can receive image information or text information from at least one robot 210_1, 210_2, or 210_3. Subsequently, the information processing system 230 can transmit information regarding the travel mode of the robots 210_1, 210_2, or 210_3 to the robots 210_1, 210_2, or 210_3.

[0039] According to one embodiment, the information processing system 230 can generate a response to a query based on images received from robots 210_1, 210_2, and 210_3, and queries related to those images, using a zero-shot-based visual language model. In this case, robots 210_1, 210_2, and 210_3 can determine a driving mode based on the response received from the information processing system 230.

[0040] Figure 3 is a block diagram showing the internal configuration of a robot 210 and an information processing system 230 according to one embodiment of the present disclosure. The robot 210 can refer to any mobile device capable of executing service applications using the robot, and capable of wired / wireless and autonomous driving, and may include, for example, robots 210_1, 210_2, 210_3 in Figure 2. As shown, the robot 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown in Figure 3, the robot 210 and the information processing system 230 may be configured to communicate information and / or data via a network 220 using their respective communication modules 316 and 336. The input / output device 320 may be configured to input information and / or data to the robot 210 via the input / output interface 318, or to output information and / or data generated by the robot 210.

[0041] The information processing system 230 may include various processors (e.g., GPU, CPU, etc.) and memory for implementing a zero-shot-based visual language model. The information processing system 230 can generate responses to images and / or text received from the robot 210 using the zero-shot-based visual language model. The information processing system 230 can transmit the generated responses to the robot 210.

[0042] The memories 312 and 332 may include any non-temporary computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include non-volatile permanent mass storage devices such as ROM (read-only memory), disk drives, SSDs (solid-state drives), and flash memory. As another example, non-volatile permanent mass storage devices such as ROM, SSDs, flash memory, and disk drives may be included in the robot 210 or the information processing system 230 as separate permanent storage devices distinct from the memories. The memories 312 and 332 may also store an operating system and at least one program code (for example, code for a robot-based service application installed and driven on the robot 210, or code for a robot control application installed and driven on the information processing system 230).

[0043] Such software components may be loaded from a computer-readable recording medium separate from memories 312 and 332. Such a computer-readable recording medium may include recording media that can be directly connected to such robot 210 and information processing system 230, and may include computer-readable recording media such as floppy disks, disks, tapes, DVD / CD-ROM drives, and memory cards. As another example, software components may be loaded into memories 312 and 332 via a communication module rather than a computer-readable recording medium. For example, at least one program may be loaded into memories 312 and 332 based on a computer program installed by a file provided via the network 220 by a developer or a file distribution system that distributes application installation files.

[0044] Processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to processors 314, 334 by memory 312, 332 or by communication modules 316, 336. For example, processors 314, 334 may be configured to execute instructions received according to program code stored in a recording device such as memory 312, 332.

[0045] Communication modules 316 and 336 can provide configurations or functions for the robot 210 and the information processing system 230 to communicate with each other via the network 220, and can provide configurations or functions for the robot 210 and / or the information processing system 230 to communicate with other robots or other systems (for example, another cloud system). For example, information generated by the robot 210's processor 314 according to program code stored in a recording device such as memory 312 (e.g., captured images, queries regarding driving modes) may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, control signals and commands provided under the control of the information processing system 230's processor 334 may be received by the robot 210 via the communication module 336 and the network 220 through the robot 210's communication module 316. For example, the robot 210 can receive responses to queries from the information processing system 230 via the communication module 316.

[0046] The input / output interface 318 may also be a means for interface with the input / output device 320. For example, the input device may include a camera including an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. In another example, the input / output interface 318 may be a means for interface with a device that integrates a configuration or function for input and output into one, such as a touchscreen. For example, when the processor 314 of the robot 210 processes instructions of a computer program loaded into memory 312, a service screen consisting of information and / or data provided by the information processing system 230 or other robots 210 may be displayed on the display via the input / output interface 318. In Figure 3, the input / output device 320 is shown not to be included in the robot 210, but is not limited to this, and may consist of the robot 210 and one device. Also, the input / output interface 338 of the information processing system 230 may be connected to the information processing system 230 or a means for interface with input or output devices (not shown) that the information processing system 230 may include. In Figure 3, the input / output interfaces 318 and 338 are shown as elements configured separately from the processors 314 and 334. However, the system is not limited to this configuration, and the input / output interfaces 318 and 338 may be configured to be included within the processors 314 and 334.

[0047] The robot 210 and the information processing system 230 may include more components than those shown in Figure 3. However, it is not necessary to explicitly show most of the conventional technical components. According to one embodiment, the robot 210 can be realized to include at least some of the input / output devices 320 described above. The robot 210 may also further include other components such as a transceiver, a GPS (Global Positioning system) module, a camera, various sensors, and a database. For example, the robot 210 may include components that a robot includes for autonomous navigation, and can be realized to further include various components such as various sensors such as an accelerometer, ultrasonic sensor, gyroscope, proximity sensor, and weight sensor, a depth camera, LiDAR, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.

[0048] According to one embodiment, the processor 314 of the robot 210 may be configured to autonomously navigate under the control of an information processing system 230. In this case, program code related to this may be loaded into the memory 312 of the robot 210. While the robot 210 is moving, the processor 314 of the robot 210 can receive information and / or data provided by the input / output device 320 via the input / output interface 318, or receive information and / or data from the information processing system 230 via the communication module 316, process the received information and / or data, and store it in the memory 312. In addition, such information and / or data can be provided to the information processing system 230 via the communication module 316.

[0049] While the robot is moving, the processor 314 can receive or receive selected text, images, videos, audio, etc., through input devices such as a touchscreen, keyboard, audio sensor, and / or image sensor, camera, and microphone connected to the input / output interface 318. The received text, images, videos, and / or audio can be stored in the memory 312 or provided to the information processing system 230 via the communication module 316 and network 220. For example, the processor 314 can receive information related to user authentication, etc., via input devices such as a touchscreen and keyboard. The received requests and / or information may be provided to the information processing system 230 via the communication module 316 and network 220.

[0050] The processor 314 of the robot 210 may be configured to manage, process, and / or store information and / or data received from the input / output device 320, other robots, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 may be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the robot 210 can output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 can also display the received information and / or data on the screen of the robot 210.

[0051] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from multiple robots 210 and / or multiple external systems (e.g., terminals used by robot 210 users). The information and / or data processed by the processor 334 can be provided to the robots 210 via the communication module 336 and the network 220. In Figure 3, the information processing system 230 is shown as a single system, but is not limited to this and may consist of multiple systems / servers.

[0052] Figure 4 shows an example of a driving mode determining factor according to one embodiment of the present disclosure. In one embodiment, various factors may exist that determine the driving mode while the robot 410 is traveling. For example, the driving mode determining factors may include the presence or absence of people 420, the presence or absence of other robots 430, whether the width of the passage is narrower than a predetermined value 440, whether the social distance between the robot and the person is smaller than a predetermined value 450, and the presence or absence of landmarks 460. Specifically, if there are people or other robots in the path of the robot 410, there is a risk of collision with the robot 410, so the driving speed needs to be adjusted. Also, if the width of the passage is narrow, the driving speed needs to be adjusted so that the robot 410 does not collide with the walls on either side. Furthermore, since people may feel anxious and / or uncomfortable when a robot is nearby, the driving speed needs to be adjusted to maintain a social distance between the robot 410 and the person. Also, if there is a specific landmark (e.g., a "slow down" sign), the driving speed of the robot 410 needs to be adjusted to comply with social rules or the rules of the robot network. Thus, depending on the driving mode determining factors, one of several driving modes can be determined. Figure 4 shows five determinants of the driving mode, but it is not limited to these, and other factors that may affect the driving mode may be added or some may be removed.

[0053] Figure 5 shows an example in which the robot's travel mode 540 is determined by one embodiment of the present disclosure. In one embodiment, the robot can use a camera to capture an image 510 and transmit it to an information processing system 230 or a server. The robot can also transmit a query 520 regarding the image 510 to the information processing system 230 or a server. Here, the query 520 may relate to factors determining the travel mode. For example, the query 520 may include questions relating to factors determining the travel mode, such as "Are there many people?", "Is there a robot?", "Is the whole space narrow?", "Is the social distance between robot and human enough?", and "Is there a 'fast' sign?".

[0054] In one embodiment, a zero-shot-based visual-language model 530 can generate a response based on an image 510 and a query 520. For example, the visual-language model 530 can generate an answer (e.g., yes / no) to each of several questions included in the query 520. Here, the visual-language model 530 may include, but is not limited to, a Visual Question Answering (VQA) model or a Large Language-and-Vision Model (LLVM).

[0055] In one embodiment, the robot can generate a response table containing responses generated by the visual language model 530. Based on the response table, the robot can determine a travel mode 540. An example of how the robot determines a travel mode based on the response table will be described in detail later with reference to Figure 6.

[0056] Figure 6 shows an example of a response table 600 according to one embodiment of the present disclosure. In one embodiment, the response table 600 may include responses to queries regarding a plurality of driving mode determinants. For example, the response "yes" may be generated for the questions "Are there many people?" and "Is there a robot?" in an image (e.g., 510 in Figure 5), and the responses "yes" or "no" may be generated for "Is the whole space narrow?", "Is the social distance between robot and human enough?", and "Is there a fast sign?". In this case, the response to each of the plurality of queries may be reflected in the response table 600.

[0057] In one embodiment, one of a plurality of pre-set driving modes can be determined based on a response table 600. Here, each of the plurality of pre-set driving modes may correspond to each of the plurality of response combinations included in the response table 600. For example, if the "person" and "robot" items are "yes" and the "narrow width", "social distance", and "landmark" items are "no" in the response table 600, the driving mode can be determined as "autonomous mode". As another example, if the "person", "narrow width", and "landmark" items are "no" and the "robot" and "social distance" items are "yes" in the response table 600, the driving mode can be determined as "strict mode". Furthermore, different weights of responses may be applied to each of the determinants of the robot's driving mode. In this case, the robot's driving mode can be determined based on a score that reflects the weights of responses to each of the determinants of the driving mode.

[0058] Figure 7 shows an example in which the robot's driving mode 740 is determined by one embodiment of the present disclosure. In one embodiment, the robot can generate a scenario 710 that describes an image based on an image taken by the robot and transmit it to an information processing system 230 or a server. For example, the robot can use a machine learning model to generate a scenario 710 in text format based on the image, such as "There are many people in the environment, but no robot is. Also, there is no specific landmark about slow sign." Here, the machine learning model may be pre-trained to take an image as input and output text that describes the image.

[0059] In one embodiment, a zero-shot-based visual language model 730 of the server can generate a response based on a scenario 710 and a query 720 received from the robot. Specifically, the query 720 generated by the robot may include a set of pre-configured driving modes (e.g., Autonomous mode) and information describing each of the driving modes (e.g., definitions and conditions for Autonomous mode). In this case, the visual language model 730 of the server can generate a response that selects the driving mode most appropriate for the scenario 710 given by the robot, based on the information describing each of the driving modes. Here, the visual language model 730 may, but is not limited to, a Large Language Model (LLM). This allows the robot to determine the driving mode 740 based on the response generated by the visual language model 730 (e.g., option 1, "Autonomous mode").

[0060] Figure 8 shows an example in which the robot's driving mode 840 is determined according to one embodiment of the present disclosure. In one embodiment, the robot can transmit images 810 taken during autonomous driving to an information processing system 230 or a server. The robot can also transmit queries 820 to the information processing system 230 or a server, which include a plurality of pre-set driving modes and information describing each of the plurality of driving modes.

[0061] In one embodiment, the server's zero-shot-based visual language model 830 can generate a response based on an image 810 and a query 820 received from the robot. Specifically, the visual language model 830 can analyze the image 810 to understand the situation. The visual language model 830 can also generate a response that selects the most appropriate driving mode for a given situation based on information describing each of several driving modes. Here, the visual language model 830 may, but is not limited to, a very large-scale visual language model (LLVM). This allows the robot to determine a driving mode 840 based on the response generated by the server's visual language model 830.

[0062] This configuration allows the robot to understand its surroundings using a zero-shot-based visual language model applicable to real-world environments, without having to learn new dynamic objects. This improves the robot's efficiency and reduces the risk of collisions by determining and switching to the appropriate driving mode depending on the situation.

[0063] Figure 9 shows an example in which the travel mode of the robot 910 is determined by one embodiment of the present disclosure. In one embodiment, the robot 910 can determine its travel mode based on absolute position information. Specifically, when the robot 910 is passing through a specific area 930 on a travel path to a destination 920, the robot 910 can autonomously travel in a preset travel mode within the specific area 930. For example, if the specific area 930 is a darkroom, even if the robot 910 takes an image with its camera, it is difficult to recognize the surrounding environment of the robot 910 in the image. In such cases, when it is difficult for the robot 910 to determine its travel mode based on a response received from the information processing system 230 or a server, the robot 910 can travel according to a preset travel mode (e.g., a strict travel mode) corresponding to the specific area 930.

[0064] In one embodiment, the robot 910 can determine its driving mode based on the characteristics of the path. Specifically, if the robot 910 is passing over an incline 940 or terrain obstacles on the path to the destination 920, the robot 910 can autonomously drive in a preset driving mode. For example, if the robot 910 is passing over an incline 940 on the path, there is a risk of the robot 910 tipping over depending on its speed. Thus, if the robot 910 cannot determine its driving mode based on a response received from the information processing system 230 or the server, the robot 910 can autonomously drive in a preset driving mode based on the incline or obstacles on the path.

[0065] In one embodiment, in sections other than the aforementioned regions 930 and 940, the robot 910 may communicate freely with the server. In this case, the robot 910 can determine its driving mode in real time based on the response received from the server.

[0066] Figure 10 is a flowchart illustrating an example of Method 1000 according to one embodiment of the present disclosure. In one embodiment, Method 1000 may be performed by at least one processor of a robot. Method 1000 may be initiated by the processor sending an image captured by the robot's camera and a query relating to the image to a server (S1010). The query may include a query relating to factors determining the driving mode. For example, the factors determining the driving mode may include at least one of the following: the presence or absence of a person, the presence or absence of a robot, the presence or absence of a landmark, whether the width of the passage is narrower than a predetermined value, or whether the social distance between the robot and the person is less than a predetermined value.

[0067] Subsequently, the processor can receive a response to the query from the server based on a zero-shot-based visual language model (S1020). The processor can then determine the robot's driving mode based on the response (S1030). Specifically, the processor can determine one of several pre-configured driving modes. Here, each of the driving modes may be configured to differ in at least one of the following: the robot's driving speed or the robot's autonomous driving tolerance level.

[0068] In one embodiment, the processor can generate a response table containing responses for each of the determinants of the travel mode based on the responses. In this case, the processor can determine one of a plurality of pre-set travel modes based on the response table. In the step of determining the travel mode of the robot, the weights of the responses for each of the determinants of the travel mode may be different.

[0069] In one embodiment, the processor can send a scenario to the server that depicts an image based on an image. In this case, the query may include information describing a set of pre-configured driving modes. The processor can also determine one of the driving modes based on the response generated based on the scenario and the query.

[0070] In one embodiment, the query may include information describing a set of pre-configured driving modes. In this case, the processor can determine one of the multiple driving modes based on the image and the response generated based on the query.

[0071] In one embodiment, the processor can determine the robot's travel mode if it receives the same response from the server a predetermined number of times. Alternatively, if the processor cannot determine the robot's travel mode based on the response, it can determine the robot's travel mode based on the robot's absolute position information. Or, if the processor cannot determine the robot's travel mode based on the response, it can determine the robot's travel mode based on the slope or obstacles on the robot's travel path.

[0072] In one embodiment, the response may be generated based on the characteristics of a particular object within the image. Alternatively, the response may be generated based on the behavior of a particular object within the image.

[0073] In one embodiment, responses may be generated for multiple different queries. That is, steps S1010 and S1020 described above may be repeated. For example, a query regarding a driving mode based on an image (or a scenario depicting an image) generated by the robot is first sent to the server, and the server can send a response to that query to the robot. Subsequently, the robot can send a subsequent query to the server regarding whether the person in the image is an infant or an elderly person, and the server can send a response to that query to the robot. The robot can also send a subsequent query to the server regarding whether the person in the image is moving or standing, and the server can send a response to that query to the robot. The robot can determine a driving mode based on responses to three different queries. The queries generated by the robot are not predetermined but are generated arbitrarily in real time, and by combining responses to multiple arbitrarily generated queries, it is possible to determine a driving mode with higher accuracy in real time.

[0074] The methods described above may be provided as computer programs stored on computer-readable recording media for execution on a computer. The media may permanently store computer-executable programs or temporarily store them for execution or download. The media may also be various recording or storage means in the form of single or multiple hardware combinations, and is not limited to media directly connected to a computer system, but may be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and media configured to store program instructions, such as ROM, RAM, and flash memory. Other examples of media include recording media and storage media managed by app stores that distribute applications, and sites and servers that supply and distribute various other software.

[0075] The methods, operations, or techniques described herein can be implemented by a variety of means. For example, such techniques can be implemented by hardware, firmware, software, or a combination thereof. A person of ordinary skill would understand that the various exemplary logic blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein can be implemented by electronic hardware, computer software, or a combination thereof. To illustrate this compatibility of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above from a functional standpoint. Whether such functions are implemented as hardware or as software depends on the design requirements imposed on the particular application and the overall system. A person of ordinary skill may implement the functions described in a variety of ways for their respective specific applications, but such implementation should not be construed as a departure from the scope of this disclosure.

[0076] In hardware implementation, the processing unit used to perform the technique may be implemented in one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, computers, or combinations thereof.

[0077] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with this disclosure may be implemented or run in any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gates and transistor logic, discrete hardware components, or any other devices designed to perform the functions described herein. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other combination of configurations.

[0078] In implementing firmware and / or software, the technique can be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), or magnetic or optical data storage devices. The instructions can be executed by one or more processors, which can then be used to perform specific aspects of the functions described herein.

[0079] When implemented as software, the above techniques may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. The computer-readable medium includes any medium that facilitates the transmission of a computer program from one location to another, and includes both computer storage media and communication media. The storage medium may be any available medium accessible by a computer. In non-limiting examples, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that is used to transport or store the desired program code in the form of instructions or data structures. Furthermore, any connection is appropriately made on the computer-readable medium.

[0080] For example, when software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair wire, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair wire, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. The terms "disk" and "disc" as used in this application include CDs, laserdiscs, optical discs, DVDs (digital versatile discs), floppy disks, and Blu-ray discs, where "disks" typically reproduce data magnetically, while "discs" reproduce data optically using a laser. The combinations considered above should also be included within the scope of computer-readable media.

[0081] The software module may be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to the processor so that the processor can read information from or record information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and storage medium may reside within an ASIC. The ASIC may reside within a user terminal. Alternatively, the processor and storage medium may exist as separate components within the user terminal.

[0082] Although the embodiments described above are described as utilizing aspects of the subject matter currently disclosed in one or more standalone computer systems, the disclosure is not limited thereto and can be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in this disclosure may be implemented with multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include PCs, network servers, and portable devices.

[0083] While this specification has described the present disclosure in relation to some embodiments, various modifications and alterations can be made without departing from the scope of the present disclosure as understandable to a person of ordinary skill in the art to which the invention of this disclosure pertains. Such modifications and alterations should be deemed to fall within the scope of the claims appended to this specification.

Claims

1. A method for determining a robot's movement mode, which is executed by at least one processor included in the robot, The steps include sending an image captured by the robot's camera and a query related to the image to a server, The steps include receiving a response to the query from the server based on a visual language model, A step of determining the robot's driving mode based on the response, A method for determining the robot's driving mode, including the method described above.

2. The robot driving mode determination method according to claim 1, wherein the query includes a query relating to factors determining the driving mode.

3. The robot driving mode determination method according to claim 2, wherein the driving mode determination factor includes at least one of the following: the presence or absence of a person, the presence or absence of a robot, the presence or absence of a landmark, whether the width of the passage is narrower than a predetermined value, or whether the social distance between the robot and the person is smaller than a predetermined value.

4. Step 1: Based on the above response, generate a response table that includes responses for each of the driving mode determining factors. It further includes, The step of determining the robot's driving mode is: The step of determining one of a set of pre-configured driving modes based on the response table. A method for determining a robot's driving mode according to claim 2, including the method described in claim 2.

5. The robot driving mode determination method according to claim 4, wherein in the step of determining the driving mode of the robot, the response weights for each of the driving mode determination factors are different.

6. A step of sending a scenario to the server that depicts the image based on the aforementioned image. The robot driving mode determination method according to claim 1, further comprising:

7. The aforementioned query includes information describing a set of pre-configured driving modes, The step of determining the robot's driving mode is: A step of determining one of the plurality of driving modes based on the response generated based on the above scenario and the above query. A method for determining a robot's driving mode according to claim 6, including the method described in claim 6.

8. The aforementioned query includes information describing a set of pre-configured driving modes, The step of determining the robot's driving mode is: A step of determining one of the plurality of driving modes based on the image and the response generated based on the query. A method for determining a robot's driving mode according to claim 1, including the method described in claim 1.

9. The step of determining the robot's movement mode is: Step to determine one of several pre-set driving modes. Includes, The robot driving mode determination method according to claim 1, wherein each of the plurality of driving modes is set to differ in at least one of the robot's driving speed or the robot's autonomous driving tolerance level.

10. The step of determining the robot's driving mode is: If the same response is received from the server a predetermined number of times, the robot's driving mode is determined. A method for determining a robot's driving mode according to claim 1, including the method described in claim 1.

11. The step of determining the robot's driving mode is: If the robot's travel mode cannot be determined based on the response, the robot's travel mode is determined based on the robot's absolute position information. A method for determining a robot's driving mode according to claim 1, including the method described in claim 1.

12. The step of determining the robot's driving mode is: If the robot's travel mode cannot be determined based on the response, the robot's travel mode is determined based on the slope or obstacles on the robot's travel path. A method for determining a robot's driving mode according to claim 1, including the method described in claim 1.

13. The robot driving mode determination method according to claim 1, wherein the response is generated based on the characteristics of a specific object in the image.

14. The robot driving mode determination method according to claim 1, wherein the response is generated based on the movement of a specific object in the image.

15. A computer-readable non-temporary recording medium that stores command words for executing the robot driving mode determination method described in claim 1 using a computer.

16. It is a robot, Communication module and Memory and The system includes at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, The at least one program is Images captured by the robot's camera and queries related to those images are sent to the server. Based on the visual language model, the response to the query is received from the server, A robot, including a command word for determining the robot's travel mode based on the response.