Information processing device, information processing method, and program
The information processing apparatus uses a large language model to interpret user instructions in natural language and write teaching information to environmental maps, addressing the challenge of interpreting and executing user instructions in existing technologies.
Patent Information
- Application Number
- PCT/JP2024/036975
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-10-17
- Publication Date
- 2025-06-19
AI Technical Summary
Existing technologies struggle to interpret user instructions in natural language and effectively write teaching information to environmental maps, making it difficult to grasp the user's intended operations on map data.
An information processing apparatus and method that utilize a large language model to call a map function corresponding to a user's language instruction, allowing the apparatus to write teaching information to an environmental map, enabling users to provide instructions in natural language for operations on the map.
Enables accurate interpretation and execution of user instructions in natural language, allowing the system to effectively write and retrieve teaching information from environmental maps, thereby improving the reproducibility and accuracy of operations performed by mobile bodies.
Smart Images

Figure JP2024036975_19062025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present disclosure relates to an information processing device, an information processing method, and a program.
[0002] In recent years, the use of natural language utterances by users as input has been considered. In such cases, it is necessary to convert the natural language expressions uttered by the user into information appropriate for subsequent processing.
[0003] For example, Patent Document 1 listed below discloses that in order to search for an object perceived by a user making an inquiry within map data, a search key for the object is identified from a natural language expression spoken by the user.
[0004] Japanese Patent Application Laid-Open No. 2023-42066
[0005] However, with the technology disclosed in Patent Document 1, although it is possible to extract search key words from the natural language spoken by the user, it is difficult to interpret the content of the user's speech and understand the user's instructions. Specifically, with the technology disclosed in Patent Document 1, it is difficult to understand the content of operations on map data instructed by the natural language spoken by the user.
[0006] Therefore, the present disclosure proposes a new and improved information processing device, information processing method, and program that are capable of writing information to an environmental map in response to a user's instruction in natural language.
[0007] According to the present disclosure, an information processing device is provided that includes a function calling unit that uses a large-scale language model to call a map function corresponding to an input user's language instruction from a predetermined group of map functions, and a function executing unit that uses the map function to write teaching information indicated by the language instruction into an environmental map.
[0008] In addition, according to the present disclosure, there is provided a computer-based information processing method, which includes using a large-scale language model to call a map function corresponding to an input user's language instruction from a predetermined group of map functions, and using the map function to write teaching information indicated by the language instruction into an environmental map.
[0009] Furthermore, according to the present disclosure, a program is provided for causing a computer to function as a function calling unit that uses a large-scale language model to call a map function corresponding to an input user's language instruction from a predetermined group of map functions, and a function executing unit that uses the map function to write teaching information indicated by the language instruction into an environmental map.
[0010] FIG. 1 is an explanatory diagram showing a first overall configuration of a system according to an embodiment of the present disclosure. FIG. 2 is an explanatory diagram showing a second overall configuration of a system according to an embodiment of the present disclosure. FIG. 3 is a block diagram showing a functional configuration of an information processing device according to an embodiment of the present disclosure. FIG. 4 is an explanatory diagram explaining a technique for generating an environmental map representing the environment around a moving object and estimating the self-position of the moving object from the position and attitude of the moving object. FIG. 5 is a table diagram showing an example of a key frame table. FIG. 6 is a table diagram showing an example of a landmark table. FIG. 7 is a flowchart diagram showing an example of an operation flow of an information processing device according to an embodiment of the present disclosure. FIG. 8 is a table diagram showing an example of a waypoint table. FIG. 9 is a table diagram showing an example of an action table. FIG. 10 is a table diagram showing an example of a landmark group table. FIG. 11 is a table diagram showing an example of a text table. FIG. 12 is an explanatory diagram showing an overview of a modified example of a system according to an embodiment of the present disclosure. FIG. 13 is an explanatory diagram showing an application example of a system according to an embodiment of the present disclosure. FIG. 14 is an explanatory diagram showing an application example of a system according to an embodiment of the present disclosure. FIG. 15 is a block diagram showing an example of a hardware configuration of an information processing device according to an embodiment of the present disclosure.
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0012] The description will be given in the following order: 1. Overall system configuration 1.1. First overall configuration 1.2. Second overall configuration 2. Configuration of information processing device 3. Operation example 4. Modification 5. Application example 6. Hardware configuration example
[0013] 1. Overall Configuration of the System First, the overall configuration of a system according to an embodiment of the present disclosure will be described with reference to FIGS. 1 and 2. FIG.
[0014] (1.1. First Overall Configuration) Fig. 1 is an explanatory diagram showing a first overall configuration of a system according to this embodiment. As shown in Fig. 1, a system 1 according to the first overall configuration includes a user terminal T, a cloud network C in which a large language model (LLM) M is implemented, and a mobile object R.
[0015] The system 1 is a system that uses a large-scale language model M to convert a linguistic instruction of a user U input to a user terminal T into a map function that writes instruction information into an environmental map held by a mobile body R. By writing the instruction information into the environmental map held by the mobile body R, the user U can instruct the mobile body R on information about the environment by giving instructions in natural language. Furthermore, the mobile body R operates based on the instruction information written into the environmental map, and can repeatedly and reproducibly execute an operation that reflects the linguistic instruction from the user U.
[0016] The user terminal T is an information terminal carried by the user U, and receives linguistic instructions using natural language from the user U. The user terminal T may be, for example, a smartphone, a smart watch, a tablet terminal, a wearable terminal, a laptop computer, or a desktop computer. For example, the user terminal T may receive linguistic instructions from the user U by collecting speech in natural language uttered by the user U using a microphone or the like and converting the collected speech into text. The user terminal T may also receive, as linguistic instructions from the user U, text in natural language input by the user U using a keyboard, a touch panel, or the like.
[0017] The cloud network C is a network including a group of computers such as servers. The cloud network C provides computational resources (e.g., the usage time of a processing unit, memory usage, etc.) for executing the large-scale language model M.
[0018] The large-scale language model M is an inference model that receives a linguistic instruction as an input and calls an appropriate function from a group of predefined functions. The large-scale language model M may be configured as a computer language model made up of a neural network having tens of millions to billions of parameters, for example. Such a large-scale language model M is trained by self-supervised learning or semi-supervised learning using a huge amount of text, and is thereby able to output an appropriate response to an input linguistic instruction using natural language.
[0019] In the system 1 according to this embodiment, when a linguistic instruction using a natural language is input, the large-scale language model M selects and outputs a function corresponding to the input linguistic instruction from a predetermined group of functions. The functions output by the large-scale language model M may include, for example, a map function that writes instruction information specified by the linguistic instruction into an environmental map held by the mobile object R. The functions output by the large-scale language model M may also include, for example, an instruction function that generates an output specified by the linguistic instruction from the sensing results of the mobile object R.
[0020] The mobile body R is a robot capable of autonomous movement. Specifically, the mobile body R may be a robot that autonomously moves by generating an environmental map of the surrounding environment and estimating its own position on the environmental map using SLAM (Simultaneous Localization and Mapping). Furthermore, the mobile body R can write instruction information instructed by a linguistic instruction from the user U into the environmental map using a map function output by the large-scale language model M. This allows the mobile body R to receive instruction on information about the environment through linguistic instructions from the user U. Since the instructed information about the environment is written into the environmental map for autonomous movement, the mobile body R can repeatedly and reproducibly execute actions that reflect the linguistic instructions from the user U.
[0021] Although the mobile object R has been described above as a robot, the technology according to the present disclosure is not limited to the above example. The mobile object R may be an automobile, a motorcycle, a bicycle, a personal mobility device, an airplane, a drone, a ship, or the like.
[0022] (1.2. Second Overall Configuration) Fig. 2 is an explanatory diagram showing a second overall configuration of a system according to this embodiment. As shown in Fig. 2, a system 2 according to the second overall configuration includes a user terminal T and a mobile object R on which a large-scale language model M is implemented.
[0023] Unlike the system 1 according to the first overall configuration, the system 2 according to the second overall configuration has a large-scale language model M implemented in the mobile object R. Therefore, in the system 2 according to the second overall configuration, the mobile object R can receive a user's linguistic instruction from the user terminal T and convert the user's linguistic instruction into a map function that writes instruction information into an environmental map using the large-scale language model M. Furthermore, the mobile object R may receive a linguistic instruction from the user U directly, without going through the user terminal T.
[0024] According to this, the system 2 according to the second overall configuration can omit communication with the cloud network C compared to the system 1 according to the first overall configuration, thereby eliminating delays that occur during communication. On the other hand, the system 1 according to the first overall configuration can execute the large-scale language model M using the vast computational resources of the cloud network C compared to the system 2 according to the second overall configuration, thereby enabling faster calculation of the large-scale language model M.
[0025] 2. Configuration of Information Processing Device Next, an information processing device according to this embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the functional configuration of the information processing device 100 according to this embodiment. The information processing device 100 executes, for example, processing in the cloud network C and mobile body R of the system 1 according to the first overall configuration, or processing in the mobile body R of the system 2 according to the second overall configuration.
[0026] As shown in FIG. 3 , the information processing device 100 includes a function calling unit 110 , a function executing unit 120 , a SLAM unit 130 , a map storage unit 140 , and a mobile object control unit 150 .
[0027] The function calling unit 110 uses the large-scale language model M to call a function corresponding to a language instruction of the user U input to the input unit 220 of the user terminal T. The input unit 220 of the user terminal T may be a microphone that picks up speech in a natural language uttered by the user U, or may be a keyboard or touch panel into which text in a natural language is input by the user U.
[0028] Specifically, the function calling unit 110 may call a map function corresponding to a language instruction of the user U input to the input unit 220 of the user terminal T from a predetermined group of map functions. The predetermined group of map functions may include a function that performs an operation or search on the environmental map. For example, the predetermined group of map functions may include a function that writes position information, annotation information, or label information specified by the language instruction of the user U onto the environmental map. Furthermore, the predetermined group of map functions may include, for example, a function that searches the environmental map for information specified by the language instruction of the user U.
[0029] Furthermore, the function calling unit 110 may call an instruction function corresponding to a linguistic instruction of the user U input to the input unit 220 of the user terminal T from a predetermined group of instruction functions. The predetermined group of instruction functions includes a function that generates an output instructed by the linguistic instruction of the user U. For example, the predetermined group of instruction functions may include a function that detects an object indicated by the linguistic instruction of the user U from the sensing result (e.g., a captured image or a depth image) of the sensor 210 mounted on the moving body R. Furthermore, the predetermined group of instruction functions may include a function that moves the moving body R by an amount and in a direction indicated by the linguistic instruction of the user U.
[0030] Furthermore, when it is necessary to execute multiple functions consecutively to execute the user U's linguistic instruction, the function calling unit 110 may call a series of multiple functions necessary to execute the user U's linguistic instruction. For example, the function calling unit 110 may call a series of multiple functions necessary to execute the user U's linguistic instruction from a predetermined map function group and a predetermined instruction function group, respectively. The subsequent function executing unit 120 can write instruction information to the environmental map instructed by the user U's linguistic instruction or generate output instructed by the user U's linguistic instruction by sequentially executing the series of multiple called functions.
[0031] The function execution unit 120 executes the function called by the function calling unit 110. By executing the called function, the function execution unit 120 can execute the content instructed by the user U in a linguistic instruction.
[0032] Specifically, the function execution unit 120 may perform an operation or search on the environmental map by executing the map function called by the function call unit 110. For example, the function execution unit 120 may write position information, annotation information, or label information specified by a linguistic instruction from the user U to the environmental map by executing the map function called by the function call unit 110. Furthermore, the function execution unit 120 may search the environmental map for information specified by a linguistic instruction from the user U by executing the map function called by the function call unit 110.
[0033] Furthermore, the function executing unit 120 may generate an output instructed by the linguistic instruction of the user U by executing the instruction function called by the function calling unit 110. For example, the function executing unit 120 may detect an object instructed by the linguistic instruction of the user U from the sensing results (e.g., captured images or depth images) of the sensor 210 mounted on the moving body R. Furthermore, the function executing unit 120 may move the moving body R by an amount and in a direction instructed by the linguistic instruction of the user U.
[0034] If a series of multiple functions is called by the function call unit 110 at a previous stage, the function execution unit 120 may execute the series of multiple functions sequentially. This allows the function execution unit 120 to use the output generated by the execution of the previous function (for example, information searched from the environmental map or information detected from the sensing result of the sensor 210) when executing the subsequent function.
[0035] The SLAM unit 130 generates an environmental map representing the environment surrounding the mobile object R and estimates the mobile object R's own position on the environmental map by SLAM using the sensing results of the sensor 210 mounted on the mobile object R. The sensor 210 may be, for example, one or more of a light detection and ranging (LiDAR), a radio detection and ranging (Radar), a time-of-flight (ToF) sensor, an ultrasonic sensor, a monocular RGB camera, or a stereo camera. The SLAM unit 130 can generate an environmental map representing the environment surrounding the mobile object R and estimate the mobile object R's own position on the environmental map by using the sensing results of the sensor 210 (e.g., environmental depth information, a depth image, or a captured image).
[0036] Specifically, the SLAM unit 130 holds the position and orientation (keyframes) of the mobile object R at a predetermined timing, and also holds the three-dimensional positions (landmarks) of feature points estimated based on the sensing results of the sensor 210 at that position and orientation. By accumulating information on these keyframes and landmarks, the SLAM unit 130 can generate an environmental map that represents the environment around the mobile object R and estimate the mobile object R's own position.
[0037] For example, FIG. 4 shows the position and orientation O of the moving body R. L 10 is an explanatory diagram illustrating a method for generating an environmental map representing the environment around the mobile object R and estimating the self-position of the mobile object R from the above.
[0038] As shown in FIG. 4, for example, the position and orientation O of the moving body R at a predetermined timing LAt this time, the SLAM unit 130 uses triangulation to calculate the position and orientation O L The image P sensed by the sensor 210 L The two-dimensional position K of the feature point LM in L The three-dimensional positions (landmarks) of the feature points LM can be estimated from the three-dimensional positions (landmarks) of the feature points LM. Therefore, the SLAM unit 130 can generate an environmental map that represents the three-dimensional structure of the environment by collecting the three-dimensional positions of the feature points LM estimated at each position and posture of the moving body R.
[0039] Furthermore, the SLAM unit 130 matches the feature points LM whose three-dimensional positions are estimated at each position and orientation with each other, thereby obtaining the position and orientation O of the moving object R that is held. L From the current position and orientation of the moving body R, R Specifically, the SLAM unit 130 first estimates the position and orientation O L The image P sensed by L and position and attitude O R The image P sensed by R By comparing with the image P L and Image P R The corresponding feature points LM on the image are matched.
[0040] Next, the SLAM unit 130 calculates the image P L Upper two-dimensional position K L The three-dimensional position of the feature point LM estimated from the image P R The estimated two-dimensional position K R The SLAM unit 130 then compares the three-dimensional position of the feature point LM estimated from the position and orientation O by solving a PnP (Perspective-n-Point) problem. L and position and posture O R Therefore, the SLAM unit 130 can calculate the relative position and orientation difference between the position and orientation O and the position and orientation O. L By adding R The absolute position and orientation of the object can be estimated.
[0041] The map storage unit 140 stores the environmental map generated by the SLAM unit 130. The map storage unit 140 may be configured, for example, with a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device.
[0042] Specifically, the map storage unit 140 may have a keyframe table and a landmark table for constructing an environmental map. The keyframe table is a table that stores the positions and orientations (keyframes) of the moving body R held at a predetermined timing. The landmark table is a table that stores the three-dimensional positions (landmarks) of feature points estimated from the sensing results of the sensor 210 at each position and orientation of the moving body R.
[0043] An example of a key frame table and an example of a landmark table will be described with reference to Figures 5A and 5B. Figure 5A is a table diagram showing an example of a key frame table. Figure 5B is a table diagram showing an example of a landmark table.
[0044] As shown in Fig. 5A, the keyframe table stores the position and posture (pose) of the moving body R, which are keyframes, in association with their identification information (ID). Also, as shown in Fig. 5B, the landmark table stores the three-dimensional positions (x, y, z) of landmark feature points in association with the identification information (ID) and the identification information (keyframe ID) of the position and posture (keyframe) of the moving body R used to estimate the three-dimensional position (x, y, z). The landmark table may further store the two-dimensional positions of the feature points in the images used to estimate the three-dimensional positions (x, y, z) of the feature points.
[0045] The map storage unit 140 may further store information written to the environmental map by the map function called by the function call unit 110. For example, the map storage unit 140 may have a table that stores location information, annotation information, or label information written to the environmental map based on a linguistic instruction from the user U.
[0046] The mobile object control unit 150 controls the drive unit 230 that drives the mobile object R. Specifically, the mobile object control unit 150 may autonomously move the mobile object R by controlling the drive unit 230 based on the environmental map generated by the SLAM unit 130 and the self-position estimated by the SLAM unit 130. Furthermore, when the function execution unit 120 executes an instruction function that instructs the mobile object R to operate, the mobile object control unit 150 may control the drive unit 230 so that the operation instructed by the instruction function is executed.
[0047] According to the above configuration, the information processing device 100 according to the present embodiment converts the linguistic instruction given by the user U into a function using the large-scale language model M, thereby making it possible to reflect the instruction in natural language of the user U in the operation of the mobile body R. Specifically, the information processing device 100 writes information corresponding to the instruction in natural language of the user U into an environmental map used for the autonomous movement of the mobile body R, thereby making it possible to reflect the instruction of the user U in the operation of the mobile body R.
[0048] In the above description, the information processing device 100 is described as simultaneously generating an environmental map and estimating the self-position of the moving object R on the environmental map using SLAM. However, the technology according to the present disclosure is not limited to this example. The information processing device 100 may estimate its own position on a pre-prepared environmental map based on the sensing results of the sensor 210. Examples of the sensor 210 used to estimate the self-position include an external sensor that senses the external environment, such as a LiDAR, radar, a time-of-flight (ToF) sensor, an ultrasonic sensor, a monocular RGB camera, or a stereo camera; an internal sensor that senses the internal state of the moving object R, such as an encoder, a gyro sensor, an acceleration sensor, or an inertial measurement unit (IMU); or a position information sensor, such as a global navigation satellite system (GNSS) sensor.
[0049] 3. Operation Example Next, the flow of operations of the information processing device 100 according to this embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the flow of operations of the information processing device 100 according to this embodiment.
[0050] 6, for example, a voice instruction is input from the user U to the input unit 220 of the user terminal T (S100). The user terminal T converts the input voice instruction into text (S102) and transmits the converted text to the function calling unit 110 (S104).
[0051] Next, the function calling unit 110 inputs the text transmitted from the user terminal T to the large-scale language model M, which outputs a function corresponding to the input text from a predetermined function group. This enables the function calling unit 110 to call a function corresponding to a linguistic instruction from the user U from a predetermined function group (S106). Note that the large-scale language model M may output the function corresponding to the input text as well as an argument of the function corresponding to the input text.
[0052] Here, if the called function is an instruction function that generates an output instructed by the user U's linguistic instruction, the function calling unit 110 outputs the called instruction function to the function executing unit 120 (S110). Furthermore, sensing information sensed by the sensor 210 of the mobile object R is output to the function executing unit 120 (S112). As a result, the function executing unit 120 executes the instruction function called by the function calling unit 110 using the sensing information from the sensor 210 (S114), thereby generating a response output of the instruction function (S116). The generated response output may be transmitted to the user terminal T, for example, and presented to the user U from the user terminal T (S118). Furthermore, the generated response output may be used when executing another function (S122).
[0053] On the other hand, if the called function is a map function that performs an operation or search on the environmental map, the function calling unit 110 outputs the called map function to the function executing unit 120 (S120). As a result, the function executing unit 120 executes the map function called by the function calling unit 110 (S122), thereby writing information to the environmental map and updating the environmental map (S124). The environmental map with the written information may be transmitted to, for example, the user terminal T and presented from the user terminal T to the user U (S126).
[0054] Furthermore, the flow of operations of the information processing device 100 according to this embodiment will be described using a more detailed example.
[0055] (First Specific Example) For example, assume that a user U inputs a voice instruction such as "This red carpet is where the meal is served" to a moving body R, which is a dog-type robot. The user terminal T converts the input voice instruction into text and transmits the converted text to the function calling unit 110.
[0056] Next, the function calling unit 110 inputs the text transmitted from the user terminal T into the large-scale language model M. As a result, the following functions corresponding to the input text are output from the large-scale language model M. The function calling unit 110 outputs the output functions to the function executing unit 120. Function: detect_object, argument: Red_carpet Function: set_waypoint, tag: dog_food
[0057] The function "detect_object" is an instruction function that detects the argument "Red_carpet" (red carpet) from an image captured by the sensor 210 of the moving body R. The function "set_waypoint" is a map function that sets the position and posture of the moving body R tagged with "dog_food" (rice) within the environmental map.
[0058] The function execution unit 120 first executes the function "detect_object" to detect an area corresponding to "Red_carpet" (red carpet) from the image captured by the sensor 210 of the moving object R. As a result, the function execution unit 120 outputs the detected area as a response output of the function "detect_object".
[0059] Next, the function execution unit 120 executes the function "set_waypoint" to set the position and orientation of the moving body R in the area corresponding to the detected "Red_carpet." The function execution unit 120 also sets the set position and orientation of the moving body R on the environmental map as a position and orientation tagged with "dog_food."
[0060] An example of a waypoint table that stores the positions and attitudes set in the environmental map by the function "set_waypoint" will be described with reference to Fig. 7. Fig. 7 is a table diagram showing an example of the waypoint table. The positions and attitudes of the moving body R stored in the waypoint table are an example of reference point information in the present disclosure.
[0061] 7 , the waypoint table stores the position and posture (pose) of the moving body R set as a waypoint in the environmental map, in association with identification information (ID) and keyframe identification information (keyframe ID). The position and posture (pose) of the moving body R stored in the waypoint table are defined, for example, as a relative position and posture from a keyframe specified by the keyframe ID. The position and posture of the moving body R defined in this way can track the keyframe, and therefore can automatically transition to an appropriate position and posture when the structure of the environmental map including the keyframe changes due to optimization or the like.
[0062] According to the first specific example, when the user U inputs, for example, "It's time for dinner," the moving body R can move to a position and posture tagged with "dog_food" (food) corresponding to the red carpet.
[0063] (Second Specific Example) For example, assume that a user U inputs a voice instruction such as "Jump at this location!" to a moving body R that is a dog-type robot. The user terminal T converts the input voice instruction into text and transmits the converted text to the function calling unit 110.
[0064] Next, the function calling unit 110 inputs the text transmitted from the user terminal T into the large-scale language model M. As a result, the following functions corresponding to the input text are output from the large-scale language model M. The function calling unit 110 outputs the output functions to the function executing unit 120. Function: set_waypoint Function: set_action
[0065] The function "set_waypoint" is a map function that sets the current position and posture of the moving body R indicated by "this location" within the environmental map. The function "set_action" is a map function that sets the moving body R to perform a specified action at a specified location.
[0066] The function execution unit 120 executes the function "set_waypoint" to set the current position and posture of the moving body R indicated by "this location" on the environmental map. In addition, the function execution unit 120 executes the function "set_action" to set the execution of an action called "jump" for the position and posture set by the function "set_waypoint".
[0067] An example of an action table that stores actions to be performed by the moving object R set by the function "set_action" will be described with reference to Fig. 8. Fig. 8 is a table diagram showing an example of the action table. The actions of the moving object R stored in the action table are an example of operation information in the present disclosure.
[0068] 8, the action table stores actions to be performed by the mobile object R at the position and attitude of a waypoint, in association with identification information (ID) and waypoint identification information (waypoint ID). This allows the mobile object R to perform the actions associated in the action table at the position and attitude set by the function "set_waypoint."
[0069] According to the second specific example, when the moving body R moves to the position and posture specified by the user U as "this location," it can perform the specified "jump" action.
[0070] (Third Specific Example) For example, assume that a user U inputs a voice instruction such as "Please explain this painting as '...'" to a mobile object R, which is a patrol robot. The user terminal T converts the input voice instruction into text and transmits the converted text to the function calling unit 110.
[0071] Next, the function calling unit 110 inputs the text transmitted from the user terminal T into the large-scale language model M. As a result, the following functions corresponding to the input text are output from the large-scale language model M. The function calling unit 110 outputs the output functions to the function executing unit 120. Function: detect_object, argument: picture Function: search_landmarks_from_bbox Function: set_landmark_group_id Function: set_text_on_landmark_group, argument: text
[0072] The function "detect_object" is an instruction function that detects a target "picture" (painting) from an image captured by the sensor 210 of the moving body R. The function "search_landmarks_from_bbox" is an instruction function that searches for landmarks included in the target area. The function "set_landmark_group_id" is an instruction function that groups multiple landmarks. The function "set_text_on_landmark_group" is a map function that sets text that serves as annotation information for multiple grouped landmarks.
[0073] The function executing unit 120 first executes the function "detect_object" to detect an area corresponding to a "picture" (painting) from the image captured by the sensor 210 of the moving object R. As a result, the function executing unit 120 outputs the area detected as a "picture" (painting) as a response output of the function "detect_object".
[0074] Next, the function execution unit 120 executes the function "search_landmarks_from_bbox" to search for landmarks included in the area detected as a "picture." Specifically, the function execution unit 120 searches for landmarks included in the area detected as a "picture" by matching feature points included in the detected area with landmarks that are the three-dimensional positions of the feature points on the environmental map.
[0075] Next, the function execution unit 120 executes the function "set_landmark_group_id" to group the retrieved landmarks and issue identification information for the collection of grouped landmarks.
[0076] An example of a landmark group table that stores each group of landmarks grouped by the function "set_landmark_group_id" will be described with reference to FIG. 9A . FIG. 9A is a table diagram showing an example of a landmark group table. As shown in FIG. 9A , the landmark group table stores the identification information (landmark ID) of a landmark and the identification information (group ID) of the group into which the landmark is grouped, in association with each other. Note that in FIG. 9A , for landmarks that are not grouped, "None" is entered as the identification information (group ID) of the corresponding group.
[0077] Furthermore, the function execution unit 120 executes the function "set_text_on_landmark_group" to set text as annotation information for a group of landmarks on the environmental map.
[0078] An example of a text table that stores text set for each landmark group will be described with reference to FIG. 9B . FIG. 9B is a table diagram showing an example of the text table. As shown in FIG. 9B , the text table stores text set for each landmark group in association with the landmark group's identification information (group ID). This associates the text "..." specified by the user U in a linguistic instruction with the landmark group included in the "painting" specified by the user U, thereby associating the text specified by the user U in a linguistic instruction with a collection of landmarks on the environmental map.
[0079] According to a third specific example, when a viewer of a painting asks for an explanation of the painting, the mobile body R can present the viewer with text of annotation information associated with landmarks contained in the painting.
[0080] 4. Modifications Next, a modification of the system according to this embodiment will be described with reference to Fig. 10. Fig. 10 is an explanatory diagram showing an overview of a modification of the system according to this embodiment.
[0081] 10 , in a modified system, an operator Op inputs a linguistic description of the real world W. The input linguistic description is converted into a function for writing information into an environmental map Em representing the real world W via a large-scale language model M.
[0082] Specifically, first, an operator Op inputs a natural language description of the characteristics of a target area in the real world W. Next, the input natural language description is converted by a large-scale language model M into a function for writing the described characteristics into an environmental map Em corresponding to the target area in the real world W. As a result, information In relating to the characteristics of the target area input by the operator Op is written into the environmental map Em corresponding to the target area.
[0083] According to this, the modified system according to the present embodiment makes it possible to write information In relating to the characteristics of the target area in the environmental map Em in advance. At this time, since the information In written in the environmental map Em is written in natural language, the user U can search for the information In using natural language. Therefore, the modified system according to the present embodiment makes it possible to write the information In in natural language in the environmental map Em, making it possible to search for the information In in the environmental map Em more directly and quickly using natural language.
[0084] 5. Application Examples The system according to this embodiment can be applied to, for example, the uses shown in Figures 11 to 13. Figures 11 to 13 are explanatory diagrams showing application examples of the system according to this embodiment.
[0085] For example, according to the system of this embodiment, as shown in FIG. 11, a user U can instruct a mobile object R, which is a housekeeping robot, in advance to check the trash can Ob1 at a specified time every day and clean up the trash in the trash can Ob1.
[0086] Furthermore, according to the system of this embodiment, as shown in FIG. 12, a user U, who is a curator, can input in advance a description and annotation of the exhibit Ob2 to a mobile body R, which is a guide robot that explains the exhibit Ob2 at an exhibition or the like.
[0087] Furthermore, according to the system of this embodiment, as shown in FIG. 13, a user U who is a worker in a factory or the like can input in advance the no-entry areas indicated by the no-entry sign Ob3 to the mobile body R, which is a cleaning robot.
[0088] As described above, the system according to this embodiment allows the user U to input instructions in natural language in advance to the mobile object R that moves around indoors autonomously.
[0089] 6. Hardware Configuration Example The hardware configuration of the information processing device 100 according to this embodiment will be described further with reference to Fig. 14. Fig. 14 is a block diagram showing an example of the hardware configuration of the information processing device 100 according to this embodiment.
[0090] The functions of the information processing device 100 according to this embodiment may be realized by cooperation between software and the hardware described below. The functions of the function call unit 110, the function execution unit 120, the SLAM unit 130, and the mobile object control unit 150 may be executed by, for example, the CPU 901. The function of the map storage unit 140 may be executed by, for example, the storage device 908.
[0091] As shown in FIG. 14, the information processing device 100 includes a CPU (Central Processing Unit) 901 , a ROM (Read Only Memory) 902 , and a RAM (Random Access Memory) 903 .
[0092] The information processing device 100 may further include a host bus 904a, a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, or a communication device 911. The information processing device 100 may have a processing circuit such as a DSP (Digital Signal Processor) or an ASIC (Application Specific Integrated Circuit) instead of or together with the CPU 901.
[0093] The CPU 901 functions as an arithmetic processing device or a control device, and controls operations within the information processing device 100 in accordance with various programs recorded in the ROM 902, the RAM 903, the storage device 908, or a removable recording medium attached to the drive 909. The ROM 902 stores programs used by the CPU 901, calculation parameters, etc. The RAM 903 temporarily stores programs used in the execution of the CPU 901, and parameters used during the execution of the programs.
[0094] The CPU 901, ROM 902, and RAM 903 are interconnected by a host bus 904a capable of high-speed data transmission. The host bus 904a is connected to an external bus 904b, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 904. The external bus 904b is connected to various components via an interface 905.
[0095] The input device 906 is a device that accepts input from a user, such as a mouse, keyboard, touch panel, button, switch, or lever. The input device 906 may also be a microphone that detects the user's voice. The input device 906 may also be, for example, a remote control device that uses infrared rays or other radio waves, or may be an externally connected device that supports operation of the information processing device 100.
[0096] The input device 906 further includes an input control circuit that outputs an input signal generated based on information input by the user to the CPU 901. By operating the input device 906, the user can input various data to the information processing device 100 or instruct the information processing device 100 to perform processing operations.
[0097] The output device 907 is a device that can visually or audibly present information acquired or generated by the information processing device 100 to a user. The output device 907 may be, for example, a display device such as an LCD (Liquid Crystal Display), a PDP (Plasma Display Panel), an OLED (Organic Light Emitting Diode) display, a hologram, or a projector, a sound output device such as a speaker or headphones, or a printing device such as a printer. The output device 907 can output information acquired by processing by the information processing device 100 as video such as text or an image, or sound such as voice or audio.
[0098] The storage device 908 is a data storage device configured as an example of a storage unit of the information processing device 100. The storage device 908 may be configured, for example, by a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 can store programs executed by the CPU 901, various data, various data acquired from the outside, and the like.
[0099] The drive 909 is a device for reading or writing data from or to a removable recording medium such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and is built into or externally attached to the information processing device 100. For example, the drive 909 can read information recorded on an attached removable recording medium and output the information to the RAM 903. The drive 909 can also write data to an attached removable recording medium.
[0100] The connection port 910 is a port for directly connecting an external device to the information processing device 100. The connection port 910 may be, for example, a Universal Serial Bus (USB) port, an IEEE 1394 port, or a Small Computer System Interface (SCSI) port. The connection port 910 may also be an RS-232C port, an optical audio terminal, or a High-Definition Multimedia Interface (HDMI) (registered trademark) port. By connecting the connection port 910 to an external device, various types of data can be transmitted and received between the information processing device 100 and the external device.
[0101] The communication device 911 is, for example, a communication interface configured with a communication device for connecting to the communication network 920. The communication device 911 may be, for example, a communication card for a wired or wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), or WUSB (Wireless USB). The communication device 911 may also be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various types of communication.
[0102] The communication device 911 can transmit and receive signals, for example, via the Internet or other communication devices using a predetermined protocol such as TCP / IP. The communication network 920 connected to the communication device 911 is a wired or wireless network, and may be, for example, an Internet communication network, a home LAN, an infrared communication network, a radio wave communication network, or a satellite communication network.
[0103] It is also possible to create a program that causes hardware such as the CPU 901, ROM 902, and RAM 903 built into a computer to perform functions equivalent to those of the information processing device 100. It is also possible to provide a computer-readable recording medium on which the program is recorded.
[0104] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0105] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0106] Note that the following configurations also fall within the technical scope of the present disclosure. (1) An information processing device comprising: a function calling unit that uses a large-scale language model to call a map function from a predetermined group of map functions that corresponds to an input linguistic instruction from a user; and a function executing unit that uses the map function to write instruction information specified by the linguistic instruction to an environmental map. (2) The information processing device according to (1), wherein the environmental map is a map generated based on sensing results of a surrounding environment of an autonomously moving mobile body. (3) The information processing device according to (2), wherein the environmental map is used for self-location estimation of the mobile body. (4) The information processing device according to (3), further comprising a mobile body control unit that controls operation of the mobile body based on the instruction information written to the environmental map. (5) The information processing device according to any one of (2) to (4), wherein the environmental map includes multiple positions and orientations of the mobile body and three-dimensional positions of feature points estimated based on sensing results of the mobile body at the multiple positions and orientations. (6) The information processing device according to (5), wherein the instruction information includes text information included in the linguistic instruction. (7) The information processing device according to (6), wherein the text information is written on the environmental map corresponding to the three-dimensional positions of the feature points. (8) The information processing device according to any one of (5) to (7), wherein the instruction information includes reference point information indicating a position and attitude of the moving object. (9) The information processing device according to (8), wherein the reference point information is further associated with action information indicating an action to be performed by the moving object at the position and attitude indicated by the reference point information. (10) The information processing device according to (8) or (9), wherein the reference point information is set based on the position and attitude of the nearest moving object included in the environmental map. (11) The information processing device according to any one of (5) to (10), wherein the function calling unit further calls an instruction function corresponding to the language instruction from a predetermined group of instruction functions, and the function executing unit uses the instruction function to generate an output indicated by the language instruction from a sensing result of the moving body.(12) The information processing device according to any one of (5) to (11), wherein the function execution unit detects the feature points from an image captured by the moving object. (13) An information processing method by a computer, comprising: using a large-scale language model to call a map function from a predetermined group of map functions corresponding to an input linguistic instruction of a user; and using the map function to write instruction information specified by the linguistic instruction into an environmental map. (14) A program for causing a computer to function as: a function calling unit that uses a large-scale language model to call a map function from a predetermined group of map functions corresponding to an input linguistic instruction of a user; and a function executing unit that uses the map function to write instruction information specified by the linguistic instruction into an environmental map.
[0107] 1, 2 System 100 Information processing device 110 Function call unit 120 Function execution unit 130 SLAM unit 140 Map storage unit 150 Mobile object control unit 210 Sensor 220 Input unit 230 Driving unit U User T User terminal R Mobile object C Cloud network M Large-scale language model
Claims
1. An information processing device comprising: a function calling unit that uses a large-scale language model to call a map function from a predetermined group of map functions corresponding to an input user's language instruction; and a function execution unit that uses the map function to write teaching information indicated by the language instruction into an environmental map.
2. The information processing device according to claim 1, wherein the environmental map is a map generated by an autonomously moving body based on the results of sensing the surrounding environment.
3. The information processing device according to claim 2, wherein the environmental map is used for self-location estimation of the moving object.
4. The information processing device according to claim 3, further comprising a mobile object control unit that controls the operation of the mobile object based on the teaching information written in the environmental map.
5. An information processing device as described in claim 2, wherein the environmental map includes multiple positions and orientations of the moving body and three-dimensional positions of feature points estimated based on sensing results of the moving body at the multiple positions and orientations.
6. The information processing device according to claim 5, wherein the instruction information includes text information included in the linguistic instruction.
7. The information processing device according to claim 6, wherein the text information is written onto the environmental map in correspondence with the three-dimensional positions of the feature points.
8. An information processing device according to claim 5, wherein the teaching information includes reference point information indicating the position and attitude of the moving body.
9. The information processing device according to claim 8, wherein the reference point information is further associated with motion information indicating a motion to be executed by the moving object at the position and posture indicated by the reference point information.
10. The information processing device according to claim 8, wherein the reference point information is set based on the position and attitude of the nearest moving object included in the environmental map.
11. The information processing device of claim 5, wherein the function calling unit further calls an instruction function corresponding to the linguistic instruction from a predetermined group of instruction functions, and the function execution unit uses the instruction function to generate an output indicated by the linguistic instruction from the sensing results of the moving body.
12. The information processing device according to claim 5, wherein the function execution unit detects the feature points from an image captured by the moving object.
13. An information processing method by a computer, comprising: using a large-scale language model to call up a map function from a predetermined group of map functions that corresponds to an input user's linguistic instruction; and using the map function to write teaching information indicated by the linguistic instruction into an environmental map.
14. A program for causing a computer to function as: a function calling unit that uses a large-scale language model to call a map function from a predetermined group of map functions that corresponds to an input linguistic instruction by a user; and a function execution unit that uses the map function to write teaching information specified by the linguistic instruction into an environmental map.
Citation Information
Patent Citations
Robot control method and device
CN117021114A
Autonomous mobile device, autonomous mobile method, and program
JP2020021257A
Contextual and user experience-based mobile robot scheduling and control
JP2021186670A
Robotic computing device with adaptive user-interaction
WO2023091160A1
Information processing device, information processing method, and program
WO2023176285A1
Cited By
LLM output result verification method and device, storage medium and equipment
CN120542582A