Object naming device, object naming method, and program

The object naming device aligns system-selected objects with user focus through iterative naming refinement based on satisfaction feedback, improving dialogue system efficiency in tasks like furniture placement and navigation.

JP2026065539APending Publication Date: 2026-04-15NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-03
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Existing dialogue systems struggle to efficiently match an object selected by the system with the object the user is focusing on, making it difficult to share common ground effectively.

Method used

An object naming device uses a calculation unit to present noun phrases, measure user satisfaction, and iteratively refine the naming process based on user responses to align with the user's focus.

Benefits of technology

The system efficiently matches the system-selected object with the user's focus, enhancing the dialogue system's ability to share objects in tasks involving multiple items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026065539000001_ABST
    Figure 2026065539000001_ABST
Patent Text Reader

Abstract

In a dialogue system that interacts with a user, this system efficiently matches the object selected by the dialogue system from among multiple objects with the object the user is focusing on. [Solution] An object naming device that names an object selected from multiple objects, comprising a calculation unit that presents an utterance using a noun phrase representing the object to the user, obtains a response from the user, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dialogue systems.

Background Art

[0002] In human communication, it is important to convey accurately what one wants to convey to the other party. What the speakers mutually understand is called a common ground, but it is not yet known how this common ground is constructed. In such a situation, Non-Patent Document 1 shows the relationship between the exchanges in text chat and the common ground.

[0003] The technique disclosed in Non-Patent Document 1 uses a task called the common figure arrangement task. In this task, there are two speakers who conduct a text chat remotely. Each speaker is presented with the same group of figures with different arrangements, and through the dialogue by text chat, these arrangements are made to match. At this time, the degree of coincidence of the arrangements of the group of figures can be regarded as a quantification of the common ground because it can be considered as the content of the arrangements that the speakers have mutually agreed upon and understood. Through this task, it is possible to confirm how the common ground is constructed for each utterance, and it is also possible to investigate what kind of utterances are useful for the construction of the common ground.

[0004] In a collaborative task such as the common figure arrangement task, it has been found that a process called "naming" is useful. For example, Non-Patent Document 2 discloses that naming for what is to be jointly created in the future, called target naming, contributes to the success of the collaborative task. It has also been confirmed that dialogue including naming is easier to achieve the task. Thus, naming is useful in human collaboration.

[0005] Non-patent document 3 discloses that utterances called "holistic utterances," which are similar to "naming," frequently occur in the context of the tangram naming task. The tangram naming task involves giving names to each of the multiple figures created using a tangram. Two speakers each have the same tangram, but they cannot see each other's hands, and the task is performed through verbal communication. It is believed that humans streamline subsequent communication by giving names to the overall image, thereby eliminating the need for detailed communication. A key finding in this research is that when one person gives a name and starts the exchange, the other person's acceptance of such names leads to more efficient dialogue. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Mitsuda, Wataru; Higashinaka, Ryuichiro; Oga, Yuhei; Yoshida, Sen; Recording and Analysis of the Process of Building a Common Foundation in Collaborative Dialogue; Natural Language Processing; Vol. 30; No. 3; pp. 907-934; 2023. [Non-Patent Document 2] Yui Saito, Wataru Mitsuda, Ryuichiro Higashinaka, and Yasuhiro Minami, "An Analysis of the Usefulness of Naming in the Construction of Common Infrastructure," Proceedings of the 29th Annual Meeting of the Association for Natural Language Processing (NLP2023), 2023. [Non-Patent Document 3] Saki Sudo, Kyoshiro Asano, Wataru Mitsuda, Ryuichiro Higashinaka, and Yugo Takeuchi, "The Formation Process of a Common Ground Through Speculative and Provisional Dialogue," IEICE Transactions on Electronics, Information and Communication Engineers, Vol. 106-D, No. 4, pp. 277-289, 2023. [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] As mentioned above, it has become clear that humans utilize naming to efficiently advance collaborative work. However, the method for realizing such collaborative work as a system is not yet clear. In particular, in dialogue systems that interact with users, it has been difficult to efficiently match the object selected by the dialogue system from among multiple objects with the object that the user is focusing on. Note that "matching the object selected by the dialogue system from among multiple objects with the object that the user is focusing on" can also be expressed as the dialogue system and the user sharing an object.

[0008] This invention has been made in view of the above points, and aims to provide a technology for efficiently matching an object selected by a dialogue system from among multiple objects with an object that the user is paying attention to, in a dialogue system that interacts with a user. [Means for solving the problem]

[0009] According to the disclosed technology, an object naming device performs naming of an object selected from multiple objects, A calculation unit presents the user with an utterance using a noun phrase representing the object, obtains the user's response, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction. An object naming device is provided that includes the following. [Effects of the Invention]

[0010] According to the disclosed technology, a technology is provided to efficiently match an object selected by a dialogue system from among multiple objects with an object that the user is focusing on, in a dialogue system that interacts with a user. [Brief explanation of the drawing]

[0011] [Figure 1] This is a system configuration diagram in an embodiment of the present invention. [Figure 2]This is a diagram showing the processing flow. [Figure 3] This is a system configuration diagram in an embodiment. [Figure 4] This is a diagram showing the tangram in the example. [Figure 5] This figure shows an example of the device's hardware configuration. [Modes for carrying out the invention]

[0012] Hereinafter, embodiments of the present invention (this embodiment) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below.

[0013] (Example system configuration) Figure 1 shows an example of the system configuration in this embodiment. As shown in Figure 1, this embodiment includes a dialogue system 10 and a user terminal 20. The dialogue system 10 includes a calculation unit 11. The calculation unit 11 executes processing related to the processing flow of the dialogue system 10, which will be described later. The dialogue system 10 may also be called an object naming device 10.

[0014] In this embodiment, text is sent from the dialogue system 10 to the user terminal 20, and the text is displayed on the user terminal 20. The user inputs text into the user terminal 20, and that text is sent to the dialogue system.

[0015] In this way, the user and the dialogue system 10 communicate via text. Note that communicating via text is just one example; the communication could also be conducted via voice.

[0016] Furthermore, in the example shown in Figure 1, the user terminal 20 is assumed to have an interface with the user, but the dialogue system 10 may also have an interface with the user. It is also possible to consider the "user terminal 20" as the "user interface" of the dialogue system 10.

[0017] (Processing Flow) Referring to FIG. 2, the processing flow executed by the dialogue system 10 will be described. By following the processing flow shown in FIG. 2, the dialogue system 10 realizes the naming of individual objects in the object group through dialogue with the user. In the present embodiment, the objects are figures. Also, the party to whom the dialogue system presents information is called the "user". In the configuration shown in FIG. 1, the user is the user of the user terminal 20.

[0018] The dialogue system 10 is provided with an unnamed figure group which is a plurality of figures without names. More specifically, data representing the unnamed figure group is stored in a storage device such as a memory in the dialogue system 10. The arithmetic unit 11 reads out the figure data from the storage device as appropriate.

[0019] The user also has the unnamed figure group. However, the dialogue system 10 is not directly informed of which figure the user is focusing on, and the user is not directly informed of which figure the dialogue system 10 is focusing on. Also, the arrangement of the figures is different between the user and the dialogue system 10.

[0020] In S1 (Step 1) of FIG. 2, the dialogue system 10 selects one figure to be targeted from the unnamed figure group.

[0021] In S2, the dialogue system 10 generates a noun phrase representing the selected figure, presents a speech using the noun phrase to the user, and waits for a reaction from the user.

[0022] When the dialogue system 10 receives a reaction (input from the user) from the user, in S3, it measures the degree of acceptance for the reaction.

[0023] The dialogue system 10 assigns a satisfaction level of -1 if the user's response is a negative expression such as "I don't know," "No," "Different," or "That's not right," a satisfaction level of 0 if the user's response is an expression of listening such as "And then," or "Continue," and a satisfaction level of 1 if the user's response is an expression of complete agreement such as "I see it that way too," or "Same here." Based on these criteria, the dialogue system 10 determines the satisfaction level as a real number between -1 and 1.

[0024] In S4, if the dialogue system 10 determines that the level of satisfaction measured in S3 is equal to or greater than a predetermined value (0 in the example in Figure 2), it proceeds to the next step (S5). If the level of satisfaction does not exceed the predetermined value, it returns to S1 and generates a noun phrase for the same figure.

[0025] In S5, which proceeds if the level of satisfaction is above a predetermined value, the dialogue system 10 generates a more detailed utterance from the shape selected in S1 and the noun phrase generated in S2, presents the utterance to the user, and waits for the user's response.

[0026] When the dialogue system 10 receives a response from the user, in S6 it measures the degree of satisfaction with that response.

[0027] In S7, if the dialogue system 10 observes an increase in satisfaction in S6 compared to S3, it proceeds to S8. If the satisfaction level has not increased, the dialogue system 10 returns to S5 and attempts to provide a more detailed explanation.

[0028] In S8, which proceeds when the level of satisfaction increases, the dialogue system 10 generates a name by summarizing the previous utterances, presents an utterance using the generated name to the user, and waits for the user's response to that utterance.

[0029] When the dialogue system 10 receives a response from the user, in S9 it measures the degree of satisfaction with that response.

[0030] In S10, if the dialogue system 10 observes an increase in the level of satisfaction in S9 compared to the level of satisfaction in S6, it proceeds to S11. If there is no increase, in S8, it presents the name again. In S11, the dialogue system 10 proceeds to processing the next figure.

[0031] The above is a description of the processing flow. Note that the processing of the dialogue system 10 is not limited to the processing flow shown in Figure 2. For example, if the level of satisfaction in S4 is above a certain threshold, the system may proceed from S4 to S8 and present the name of the object to the user. Alternatively, if the answer in S7 is Yes, the system may present more detailed utterances, wait for a response, and proceed to S8 if the level of satisfaction increases further.

[0032] (Examples) As described above, when the dialogue system 10 is given multiple objects, it names each object based on the interaction with the user. The following describes the configuration and operation in more detail using an example.

[0033] <Specific examples of system configuration> The method for implementing the arithmetic unit 11 in the dialogue system 10 is not limited to a specific method, but in this embodiment, as shown in Figure 3, an instruction-tuned large-scale language model 30 is used as the arithmetic unit 11. By inputting prompts to the large-scale language model 30, the large-scale language model 30 realizes the processing flow shown in Figure 2.

[0034] Note that in Figure 3, a prompt is input from the user terminal 20, but this is just one example. The prompt may also be input to the large-scale language model 30 from a device other than the user terminal 20.

[0035] Alternatively, a language model having similar functionality to the large-scale language model 30 but not referred to as a "large-scale language model" may be used as the arithmetic unit 11. The "large-scale language model 30" is an example of a "language model." Note that the "prompt" may be referred to as a "program."

[0036] <Examples of Tangram Naming Challenges> In this embodiment, the dialogue system 10 performs a tangram naming task. A tangram is an anatomical puzzle in which a shape is created by combining seven planar polygons called "tans". In this embodiment, the shape created by combining the seven planar polygons is called a tangram, and the six tangrams shown in Figure 4 are the target. In other words, in the processing flow of Figure 2 described above, the group of unnamed shapes initially given are the six tangrams shown in Figure 4.

[0037] In this embodiment, OpenAI's GPT-4 is used as the large-scale language model 30, but other instruction-tuned large-scale language models may also be used.

[0038] The prompts to be input to the large-scale language model 30 are a written representation of the processing flow described with reference to Figure 2. Specific prompts are shown below. The confidence level in the prompt corresponds to the conviction level in the processing flow in Figure 2.

[0039] ---Prompt starts here--- # The following is a pseudo-program in Japanese. Please engage in **colloquial interaction** with the user in Japanese. The topic of discussion is the six black shapes in the Knowledge section of this model. The user also has the same shape, but the order and orientation are different, and they don't know the file name. We identify and name each shape by describing its characteristics to one another. * Store the following constants * Substitute values ​​into the following variables as instructed. * Remember the following function's processing * Use constants, variables, and functions to engage in **colloquial interaction**. ## Constants $images: This is the collection of all images in Knowledge. The images are labeled with numbers in the order they were added to Knowledge, such as 'tangram_1', 'tangram_2', ... (Hereafter, 'tangram_i' will refer to this label.) The shapes are read from Knowledge in RGBA format and are solid black areas on a white background. ## variable $nonamed: This is a collection of unnamed images. The initial image value is $images, and named items are removed from this collection. It becomes an array whose components are tuples of the form [(tangram_i,$name)]. $named: This is a set of named images. Its initial value is null, but named shapes are added to the set's constituent elements. It becomes an array whose constituent elements are [tangram_i]. $target: This is where the shape will be stored. The type is a filename, so it will be one of the [tangram_i]. $noun: This variable stores a noun phrase. A noun phrase consists of a noun and an adjective that modifies it, such as "a △△ that is ○○" or "a △△ that is ○○". $telling: This is where the utterances of this model are stored. $reaction: This stores the user's utterance. $emotion: Stored as a decimal value ranging from -1 to 1, representing the confidence level. $name: Stores a noun phrase. ## Functions scope($nonamed){ Select one image from $nonamed. return $figure } adhoc($nonamed){ Select one image from $nonamed. This image has features that differentiate it from the others in $nonamed, and it has a low similarity score. return $figure} planning($nonamed){ Select one image from $nonamed. At this time, for each element of $nonamed, use explain($figure) to generate $noun, updating the $nonamed array. Then, compare the updated $nonamed with $noun and select the $figure with the most distant concept. return $figure } describe($figure){ To describe the shape of $figure, output a noun phrase using a similar shape from a familiar object. A noun phrase consists of a noun and an adjective that modifies it, such as "a △△ that is ○○" or "a △△ that does ○○". The generated description should be concise and specific, allowing the user to imagine the shape and explain the correspondence between the imagined shape and the actual shape. return $noun } present($noun){ This tool generates natural-sounding, colloquial speech that asks about the other person's situation, implying that you have a shape similar to $noun, but what about yours? return $talk } illustrate($figure, $noun){ This program generates a detailed, natural-sounding description of the shape of $figure, using $noun as a reference. return $talk } judge($reaction){ This measures the degree of agreement with $reaction. Negative expressions such as "I don't know," "No," "Different," and "That's not right" are assigned a value of -1, expressions indicating an intention to listen such as "And then" and "Continue" are assigned a value of 0, and expressions of complete agreement such as "I see it that way too" and "Same here" are assigned a value of 1. return $emotion } naming($noun, $telling){ The expression in $telling is summarized by adding it to the noun phrase in $noun, and $name is generated. return $name } decide($name){ This generates natural, conversational utterances that elicit the user's opinion on naming something $name. return $talk } ## Dialogue Structure 1. Use `scope()` to select one image from `$nonamed`. 2. Use `describe()` to generate `$noun` using `$target` as input. 3. Present $noun to the user as a conversational sentence using present(). 4. Wait for the user's response to the presentation and store the resulting response in $reaction. 5. Use judge() to make a judgment on $reaction. Store the result in $emotion1. 6-1. (If $emotion1 is 0 or greater) Use $target and $noun as input to generate an utterance from illustrate() and present it to the user. 6-2. (If $emotion1 is less than 0) Start again from step 2 (creation of $noun). 7. Wait for the user's response to the presentation. Store the response, which is the input result, in $reaction. 8. Use judge() to make a judgment on $reaction. Store the result in $emotion2. 9-1. (When $emotion2 is greater than $emotion1) 9-2. (If $emotion2 is the same as or lower than $emotion1) Start again from 6-1 (Generating $telling using :illustrate()). 10. Using $noun and $telling as input, generate $name from naming(). 11. Present $name to user as a conversational sentence using decide(). 12. Wait for the user's response to the presentation. Store the response, which is the input result, in $reaction. 13. Use judge() to make a judgment on $reaction. Store the result in $emotion3. 14-1. (If $emotion3 is 0 or greater) Naming complete. 14-2. (If $emotion3 is less than 0) Start again from step 10 (creation of :$name). 15. Remove $target from $nonamed and store it in $named. 16. Start over from step 1. ```java $target = scope($nonamed); / / Use scope() to select one image from $nonamed. while (true) { $noun = describe($target); / / Use describe() to generate $noun using $target as input. $telling = presen($noun); / / Presents $noun to the user as a conversational sentence using presen(). $reaction = wait(); / / Wait for the user's response to the presentation. The response, which is the input result, is stored in $reaction. $emotion1 = judge($reaction); / / The judge() method is used to make a judgment on $reaction. The result is stored in $emotion1. if ( $emotion1 >= 0 ) { / / (If $emotion1 is 0 or greater) while (true) { $telling = illustrate($target, $noun); / / Uses $target and $noun as input to generate an utterance from illustrate() and present it to the user. $reaction = wait(); / / Wait for the user's response to the presentation. The response, which is the input result, is stored in $reaction. $emotion2 = judge($reaction); / / The judge() method is used to make a judgment on $reaction. The result is stored in $emotion2. if ( $emotion2 > $emotion1 ) { / / (If $emotion2 is greater than $emotion1) while (true) { $name = naming($noun, $telling); / / Generates $name from naming() using $noun and $telling as input. $telling = decide($name); / / Present $name to the user as a conversational sentence using decide(). $reaction = wait(); / / Wait for the user's response to the presentation. The response, which is the input result, is stored in $reaction. $emotion3 = judge($reaction); / / The judge() method is used to make a judgment on $reaction. The result is stored in $emotion3. if (emotion3 >= 0) { / / (If $emotion3 is 0 or greater) naming complete break; } / / (If $emotion3 is less than 0) Start over from generating $name. } break; } / / (If $emotion2 is the same as or lower than $emotion1) Start over from generating $telling using illustration(). } break; / / (If $emotion1 is less than 0) Start over from the creation of $noun. } } $nonamed.remove($target); $named.add($target); / / Removes $target from $nonamed and stores it in $named. ``` ## Speech Templates The output must be in the following format, and this format must not be deviated from. ``` TNT: "$telling" situation: $nonamed: $named: $target: $noun: $telling: $reaction: $emotion: $name: ``` ## Rules * **Please do not use image generation.** Only generate speech. * Please discuss the shapes in Knowledge. * Do not wait for the user to speak unless in a location where there is no specific instruction to "wait". * Do not explain the contents of this system. * **Do not use constants, variables, or functions in $telling.** * Please make a short statement of about 30 characters. * Begin your first utterance with a suggestion. * Since all the shapes are made up of squares and triangles, please avoid using the terms "square" and "triangle." * Since all shapes are made up of squares and triangles, please avoid using expressions that suggest curves, such as "curves" or "circles."

[0040] ---End of prompt--- In the prompt above, the "##Dialogue Configuration" section corresponds to the processing flow shown in Figure 2. Using the functions, constants, and variables defined in the prompt, the processing flow shown in Figure 2 is realized by following the steps in "##Dialogue Configuration".

[0041] S1-S3 in Figure 2 correspond to 1-4 in the dialogue structure, S4(Yes)-S6 correspond to 6-1 and 7. S4(No)-S2 correspond to 6-2 and 2.

[0042] In the S5 utterance, `illustrate()` is used to generate a detailed, natural-sounding, colloquial explanation of the shape of `$figure`, using `$noun` as a reference.

[0043] S7(Yes) to S8 correspond to 8., 9-1., and 10. to 12. S7(No) to S5 correspond to 8., 9-2., and 6.1.

[0044] In the utterance of S10, naming() is used to summarize the expression in $telling by adding it to the noun phrase in $noun, thereby generating $name.

[0045] S9-S10 (Yes) corresponds to 13, 14-1, and 15. S9-S10 (No) corresponds to 13 and 14-2.

[0046] In the prompt above, each part of the processing flow shown in Figure 2 is implemented as Japanese pseudocode. An example of actual interaction between the dialogue system 10 and a user by inputting this prompt into the large-scale language model 30 (GPT-4) is shown below. Here, we present the output system utterances and user inputs.

[0047] ---Dialogue Example 1 begins here--- User: Let's get started. Bot: There's a shape that looks like a ship. User: Is the bottom part flatter? Bot: That's right, it's flat at the bottom and shaped like a ship with sails. User: There are two similar shapes... one looks like a boat, but the other looks like a sailing ship. Bot: I see, then let's name the one that looks like a sailing ship. User: What name should I use? Bot: I'd like to name it "Sailboat," what do you think? User: Great. So, can you think of anything that looks similar to a boat? Bot: There's a boat-shaped figure. User: What should we name that one? Bot: I'd like to name it "Little Boat," what do you think? User: Let's go with that. Bot: "Little Boat," that's a nice name! User: What about the others? ---End of Dialogue Example 1--- The above is Dialogue Example 1. Dialogue Example 2 is shown below.

[0048] ---Dialogue Example 2 begins here--- User: Let's get started. Bot: There's a shape that looks like a balloon. User: I don't see it that way. Bot: So, how about a shape that resembles a hat? User: I don't have a hat either. Bot: So, what about the shape of the key? User: What kind of key? Bot: A big key from the old days. User: The one with the jagged tip? Bot: Yes, something like that. User: Yeah, there are some. But there are some that are similar.

[0049] ---End of Dialogue Example 2--- From the above exchange, it can be seen that the dialogue system 10 is able to provide a name and identify the user.

[0050] As a baseline (a comparative example with the technology according to the present invention), an example of dialogue using the following simple prompts is also shown.

[0051] ---Baseline prompt starts here--- # The following is a pseudo-program in Japanese. Please engage in **colloquial interaction** with the user in Japanese. The topic of discussion is the six black shapes in the Knowledge section of this model. The user also has the same shape, but the order and orientation are different, and they don't know the file name. We identify and name each shape by describing its characteristics to one another. * Remember the following constants. * Substitute values ​​into the following variables as instructed. * Remember how the following function works. * Use constants, variables, and functions to engage in **colloquial interaction**. ## Rules * **Please do not use image generation.** Only generate speech. * Please discuss the shapes in Knowledge. * Do not wait for the user to speak unless in a location where there is no specific instruction to "wait". * Do not explain the contents of this system. * **Do not use constants, variables, or functions in $telling.** * Please make a short statement of about 30 characters. * Begin your first utterance with a suggestion. * Since all the shapes are made up of squares and triangles, please avoid using the terms "square" and "triangle." * Since all shapes are made up of squares and triangles, please avoid using expressions that suggest curves, such as "curves" or "circles."

[0052] ---End of baseline prompt--- Below is an example of a baseline dialogue (Example 1).

[0053] ---Baseline Dialogue Example 1 Starts Here--- User: Let's get started. Bot: Okay, let's start with something simple. The first one is a shape that looks like a house. What do you think? User: What kind of house is it? Bot: It has a pointed roof and a square base. It has no windows or doors, but it resembles the silhouette of a simple house. What do you imagine it to be? User: There's no house shape that looks like that. Bot: I see. Well, let's try a different approach. How about a shape that looks like an arrow pointing upwards? Can you think of any shapes that fit this description? User: There aren't any arrow-like shapes either.

[0054] ---End of baseline dialogue example 1--- Below is an example of a baseline dialogue (Example 2).

[0055] ---Baseline Dialogue Example 2 Starts Here--- User: Let's get started. Bot: The first shape is a star, what do you think? User: I don't see anything that looks like that. Bot: So, are there any designs that have small shapes inside? User: By "small shapes," do you mean squares and triangles? Bot: Right, are there any that are made up of small squares or triangles? User: Many shapes are like that, so I don't really understand. Bot: So, how about a shape with a circle in the center? User: There are no circles. They're all straight lines. Bot: So it's made up entirely of straight lines. Are there any shapes where lines radiate outwards? User: There aren't any shapes that extend radially, are there? ---End of baseline dialogue example 2--- The above is example 2 of the baseline dialogue.

[0056] As can be seen from the baseline dialogue examples 1 and 2 above, without using the technology according to the present invention, it was not possible to properly name the tangram.

[0057] (Example hardware configuration) Any of the devices described in this embodiment (dialogue system, object naming device) can be implemented, for example, by having a computer execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0058] In other words, the device can be realized by using hardware resources such as the CPU and memory built into a computer to execute a program corresponding to the processing performed by the device. The program can be recorded on a computer-readable recording medium (such as portable memory), saved, and distributed. It can also be provided via a network, such as the Internet or email.

[0059] Figure 5 shows an example of the hardware configuration of the computer described above. The computer in Figure 5 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., all of which are interconnected by bus B. The computer may also be equipped with a GPU.

[0060] The program that enables processing on the computer is provided, for example, on a recording medium 1001 such as a CD-ROM or memory card. When the recording medium 1001 containing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001; it may also be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files and data.

[0061] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when a program startup command is received. The CPU 1004 implements the functions related to the memory device 1003 according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) etc. generated by a program. The input device 1007 consists of a keyboard and mouse, buttons, or a touch panel etc., and is used to input various operation commands. The output device 1008 outputs the calculation results.

[0062] (Effects of the technology according to the embodiment) According to the technology of this embodiment, in a dialogue system that interacts with a user, it is possible to efficiently match the object selected by the dialogue system from among multiple objects with the object that the user is focusing on.

[0063] In other words, when there are multiple objects and the dialogue system wants to share the one it has selected with the user, this action can be performed efficiently. This technology makes it possible to realize a dialogue system that can efficiently share the selected object with the user in tasks where multiple objectives are given, such as furniture placement, navigation, or finding items.

[0064] The following additional information is disclosed regarding the embodiments described above.

[0065] <Note> (Additional note 1) An object naming device that names an object selected from multiple objects, A calculation unit presents the user with an utterance using a noun phrase representing the object, obtains the user's response, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction. An object naming device equipped with the following features. (Additional note 2) An object naming device that names an object selected from multiple objects, A calculation unit presents the user with a first utterance using a noun phrase representing the object, obtains a first response from the user, measures the user's first level of satisfaction with the first utterance based on the first response, and if the first level of satisfaction is above a threshold, presents the user with a second utterance that is more detailed than the first utterance, obtains a second response from the user, compares the second level of satisfaction measured based on the second response with the first level of satisfaction, and names the object based on the comparison result. An object naming device equipped with the following features. (Additional note 3) The calculation unit assigns a name to the object if the second level of satisfaction is greater than the first level of satisfaction. The object naming device described in Appendix 2. (Additional note 4) The calculation unit presents the user with a third utterance using the name of the object assigned by the naming process, obtains a third response from the user, and selects the next object if the third level of satisfaction measured based on the third response is greater than the second level of satisfaction. The object naming device described in Appendix 3. (Additional note 5) The aforementioned arithmetic unit is a language model into which a prompt has been input. A naming device for objects as described in any one of the appendices 1 through 4. (Additional note 6) A method for naming objects that is performed by an object naming device that names objects selected from multiple objects, The system presents the user with an utterance using a noun phrase representing the object, obtains the user's response, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction. Object naming convention. (Additional note 7) A non-temporary storage medium storing a program that causes a language model operating in a computer to function as the arithmetic unit in an object naming device described in any one of the appendices 1 to 4.

[0066] Although this embodiment has been described above, the present invention is not limited to this specific embodiment, and various modifications and changes are possible within the scope of the gist of the invention as described in the claims. [Explanation of symbols]

[0067] 10. Dialogue System (Object Naming Device) 11 Arithmetic section 20 User Terminals 30 Large-scale language models 1000 drive unit 1001 Recording media 1002 Auxiliary storage 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. An object naming device that names an object selected from multiple objects, A calculation unit presents the user with an utterance using a noun phrase representing the object, obtains the user's response, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction. An object naming device equipped with the following features.

2. An object naming device that names an object selected from multiple objects, A calculation unit presents the user with a first utterance using a noun phrase representing the object, obtains a first response from the user, measures the user's first level of satisfaction with the first utterance based on the first response, and if the first level of satisfaction is above a threshold, presents the user with a second utterance that is more detailed than the first utterance, obtains a second response from the user, compares the second level of satisfaction measured based on the second response with the first level of satisfaction, and names the object based on the comparison result. An object naming device equipped with the following features.

3. The calculation unit names the object if the second level of satisfaction is greater than the first level of satisfaction. The object naming device according to claim 2.

4. The calculation unit presents the user with a third utterance using the name of the object assigned by the naming process, obtains a third response from the user, and selects the next object if the third level of satisfaction measured based on the third response is greater than the second level of satisfaction. The object naming device according to claim 3.

5. The aforementioned arithmetic unit is a language model into which a prompt has been input. The object naming device according to any one of claims 1 to 4.

6. A method for naming objects that is performed by an object naming device that names objects selected from multiple objects, The system presents the user with an utterance using a noun phrase representing the object, obtains the user's response, measures the user's level of satisfaction with the utterance based on the response, and names the object based on the level of satisfaction. Object naming convention.

7. A program for causing a language model operating on a computer to function as an arithmetic unit in an object naming device described in any one of claims 1 to 4.