Dialogue systems, dialogue methods, and programs
The dialogue system automatically names graphic objects to facilitate efficient communication between humans and systems by rearranging graphic objects based on user interactions, addressing the limitations of existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2024-10-25
- Publication Date
- 2026-05-13
AI Technical Summary
Existing dialogue systems fail to enable efficient communication between humans and systems by automatically naming graphic objects during collaborative tasks.
A dialogue system that generates graphic placement coordinates by interacting with a user, comprising a communication unit, graphic movement detection unit, and graphic movement unit to automatically name and rearrange graphic objects based on user utterances and a language model.
Enables efficient communication between users and dialogue systems by allowing automatic naming of graphic objects, similar to human communication.
Smart Images

Figure 2026077311000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for generating a target graphic arrangement by rearranging graphic objects.
Background Art
[0002] In recent years, in a dialogue system, a human can communicate with a computer such as a smart speaker and obtain various information or satisfy desires.
[0003] Also, in human-to-human communication, it is important to convey accurately what one wants to convey to the other party. The content mutually understood by the speakers is called a common ground, but how this common ground is constructed is not yet understood. In such a situation, a technique showing the relationship between the exchanges in text chat and the common ground has been disclosed (Non-Patent Document 1).
[0004] In this Non-Patent Document 1, a problem called a joint graphic arrangement problem is used. In this problem, two speakers (users) conduct a text chat. Each speaker is presented with the same group of graphic objects having different arrangements. Then, each speaker independently aligns the arrangements of each graphic object on each PC or the like through the dialogue by text chat. At this time, the degree of coincidence of the arrangements of the group of graphic objects can be regarded as a quantification of the common ground because it can be considered as the arrangement content that the speakers have mutually understood and agreed upon. Through this joint graphic arrangement problem, it is possible to confirm how the common ground is constructed for each utterance, and it is also possible to investigate what kind of utterances are useful for the construction of the common ground.
[0005] In collaborative tasks involving human interaction, such as collaborative geometric arrangement tasks, a process called "naming" has proven useful. For example, Non-Patent Literature 2 discloses that goal naming, such as "house" or "city," which represents the geometric arrangement to be created collaboratively, is likely to contribute to the success of collaborative work. It has also been confirmed that dialogues that include naming are more likely to achieve the task. Thus, naming geometric objects and geometric arrangements is useful in collaborative work involving human interaction. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Mitsuda, Wataru; Higashinaka, Ryuichiro; Oga, Yuhei; Yoshida, Sen; Recording and Analysis of the Process of Building a Common Foundation in Collaborative Dialogue; Natural Language Processing; Vol. 30; No. 3; pp. 907-934; 2023. [Non-Patent Document 2] Yui Saito, Wataru Mitsuda, Ryuichiro Higashinaka, and Yasuhiro Minami, "An Analysis of the Usefulness of Naming in the Construction of Common Infrastructure," Proceedings of the 29th Annual Meeting of the Association for Natural Language Processing (NLP2023), pp. 1985-1989, 2023. [Overview of the project] [Problems that the invention aims to solve]
[0007] However, Non-Patent Documents 1 and 2 do not allow naming processes to be performed when one party in the dialogue is a human and the other is a system.
[0008] The present invention has been made in view of the above-mentioned points, and aims to enable a dialogue system to automatically name graphic objects, thereby realizing efficient communication between a user (human) and a dialogue system, similar to communication between humans. [Means for solving the problem]
[0009] To solve the above problems, this disclosure provides a dialogue system that generates graphic placement coordinates by arranging graphic objects while interacting with a user operating a user terminal, the dialogue system comprising: a communication unit that receives user utterances sent from the user terminal; a graphic movement detection unit that uses a language model to obtain predetermined naming information for a predetermined graphic object and destination coordinates for the ID of the predetermined graphic object from the user utterances, and reads the ID of the graphic object corresponding to the predetermined naming information from a naming information management unit that manages the IDs and naming information of each graphic object; and a graphic movement unit that updates the position coordinates corresponding to the ID of the predetermined graphic object to the destination coordinates in the coordinate information management unit that manages the IDs and position coordinates of each graphic object. [Effects of the Invention]
[0010] As explained above, the present invention enables the dialogue system to automatically name graphic objects, thus achieving efficient communication between the user (human) and the dialogue system, similar to communication between humans. [Brief explanation of the drawing]
[0011] [Figure 1] This is an overall diagram of the communication system. [Figure 2] This is an electrical hardware configuration diagram of the user terminal and the dialogue system. [Figure 3] This is a functional configuration diagram of the dialogue system. [Figure 4] This is a flowchart showing the processing of the dialogue system. [Figure 5] This is a flowchart showing the processing of the dialogue system. [Figure 6] This figure shows an example of a function calling description in the shape movement detection unit. [Figure 7] This figure shows an example of a prompt from the dialogue content generation unit. [Figure 8]It is a diagram showing an example of a prompt of a dialogue content generation unit. [Figure 9] As an experimental result, it is a diagram showing an example of a dialogue between a user and a dialogue system.
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described based on the drawings.
[0013] 〔Outline of Embodiment〕 First, the outline of this embodiment will be described using FIG. 1. FIG. 1 is an overall configuration diagram of this embodiment.
[0014] As shown in FIG. 1, the communication system 10 includes a dialogue system 20 and a user terminal 80. The dialogue system 20 and the user terminal 80 can perform data communication via a communication network 100 such as the Internet.
[0015] The user terminal 80 is a notebook PC, a desktop PC, a smartphone, a tablet terminal, or the like. The user operates a mouse or the like of the user terminal 80 to move a pre-prepared graphic object (such as a triangular graphic, a quadrilateral graphic, etc.) while speaking the content of the movement (for example, "move the small triangle to the lower left") for creating the target graphic arrangement coordinates.
[0016] Hereinafter, "graphic arrangement coordinates" will be expressed as "arrangement", and "graphic object" will be expressed as "graphic".
[0017] Also, the uttered content may be input as text such as in a chat, or may be spoken orally and converted into text by the voice recognition function of the user terminal 80. The text data indicating this uttered content is transmitted to the dialogue system 20.
[0018] The dialogue system 20 is a computer. The dialogue system 20 independently creates a target layout by moving a graphic in the memory (such as the coordinate information management unit 21 described later) based on the user's utterance content sent from the user terminal 80 and the system utterance content automatically generated by the dialogue system 20, independently of the user terminal 80. Also, the dialogue system 20 transmits the data of the system utterance content it created to the user terminal 80. As a result, on the user terminal 80, the text or voice of the system utterance content is output, and the user can communicate with the dialogue system 20.
[0019] 〔Hardware Configuration〕 Next, the hardware configuration of the user terminal 80 will be described using FIG. 2. FIG. 2 is a hardware configuration diagram of the user terminal according to the embodiment.
[0020] As shown in FIG. 2, the user terminal 80 has a processor 1001, a memory 1002, an auxiliary storage device 1003, a communication device 1004, and a connection device 1005. Also, the user terminal 80 has an operation device 1006, a display device 1007, a voice input device 1008, and a voice output device 1009. Each hardware component constituting the user terminal 80 is interconnected via a bus 1010 such as a data bus.
[0021] The processor 1001 serves as a control unit that controls the entire user terminal 80 and has various arithmetic devices such as a CPU (Central Processing Unit). The processor 1001 reads and executes various programs on the memory 1002. Note that the processor 1001 may include a GPU (Graphics Processing Unit).
[0022] Memory 1002 has main memory devices such as ROM (Read Only Memory) and RAM (Random Access Memory). The processor 1001 and memory 1002 form a so-called computer, and the computer realizes various functions by having the processor 1001 execute various programs read into memory 1002.
[0023] The auxiliary storage device 1003 stores various programs and various information used when these programs are executed by the processor 1001.
[0024] The communication device 1004 is a communication device for sending and receiving various types of information with other devices (including equipment, servers, and systems).
[0025] The connection device 1005 is a connection device used to connect various sensors, external memory, etc., to the user terminal 80.
[0026] The operating device 1006 accepts input of various types of information, such as text and images, through user operation.
[0027] The display device 1007 is, for example, a device (such as a display) that displays various information acquired from the operating device 1006 or other devices as an image.
[0028] The voice input device 1008 detects voice information such as the user's voice.
[0029] The audio output device 1009 is, for example, a device that outputs various information received from other devices as audio.
[0030] Note that the dialogue system 20 has the same configuration as shown in Figure 2, so its explanation is omitted. The user terminal 80 and the dialogue system 20 do not necessarily have a voice input device 1008 and a voice output device 1009.
[0031] [Functional Configuration] Next, the functional configuration of the dialogue system according to the embodiment will be explained using Figure 3. Figure 3 is a functional configuration diagram of the dialogue system.
[0032] As shown in Figure 3, the dialogue system 20 includes a communication unit 31, a shape movement detection unit 41, a shape movement unit 42, a naming detection unit 51, a shape selection unit 52, a target naming detection unit 61, a target shape placement creation unit 62, a moving shape determination unit 63, and a speech content generation unit 71. Each of these units is a function realized by instructions from the processor 1001 in Figure 2 based on a program. In addition, the memory 1002 or auxiliary storage device 1003 of the dialogue system 20 contains a coordinate information management unit 21, a probability structure management unit 22, a naming information management unit 23, a target information management unit 24, a naming setting information management unit 25, and a speech history management unit 26.
[0033] <Each management department> The coordinate information management unit 21 manages the relationship between the ID of each shape (referred to as "shape ID") and the position (placement) coordinates of each shape. This allows the dialogue system 20 to recognize the position of each shape in the virtual space, as shown in Figure 1. Note that the ID (Identifier) is identification information.
[0034] The probability structure management unit 22 manages the probability structure P (figure|naming) (see Non-Patent Literature 2).
[0035] Here, "probability structure" refers to statistical information obtained from corpus data that already includes a correspondence between shapes and naming information, representing which shapes (shape IDs) the naming information could represent.
[0036] The naming information management unit 23 manages the association between the shape ID of each shape and its naming information.
[0037] The target information management unit 24 manages information on the target design drawings (layouts), which will be described later.
[0038] The naming setting information management unit 25 manages the naming setting information described below. The naming setting information is information that associates the figure (figure ID) determined by the moving figure determination unit 63 with the selected naming information.
[0039] The speech history management unit 26 manages the text data, which is the history of user utterances and system utterances.
[0040] <Shape movement detection unit> The shape movement detection unit 41 uses an instruction-tuned large-scale language model (hereinafter referred to as "language model M") such as GPT (Generative Pre-trained Transformer)-4 to obtain the naming information of a shape and the destination coordinates of this shape from the utterance content (text) of the user or dialogue system 20. The shape movement detection unit 41 also searches the naming information management DB 23 using the naming information obtained from language model M as a search key and identifies the shape by reading the corresponding shape ID. The shape movement detection unit 41 then outputs the shape ID and destination coordinates to the shape movement unit 42.
[0041] <Shape Movement Section> The shape movement unit 42 moves the shape associated with the shape ID to the destination coordinates based on the shape ID and destination coordinates obtained from the shape movement detection unit 41. For example, if there are multiple "rectangle" shapes, the shape movement unit 42 randomly selects a predetermined shape from among the multiple shapes with equal probability according to a uniform distribution and moves this selected shape. The shape movement unit 42 then updates the destination coordinates in the coordinate information management unit 21 with the coordinates corresponding to the shape ID of the moved shape. If the shape movement detection unit 41 cannot identify the shape from the naming information, the shape movement detection unit 41 does not perform any updates to the coordinate information management unit 21.
[0042] <Naming detection unit> The naming detection unit 51 uses the language model M to determine whether or not naming information is included in the utterance (text). If it is included, it detects (extracts) the naming information from the utterance and outputs this naming information to the shape selection unit 52. For example, the naming detection unit 51 extracts the naming information ("eye") from the utterance ("Let's call it an eye").
[0043] <Shape Selection Area> The shape selection unit 52 uses the probability structure P(shape|naming), which is managed by the probability structure management unit 22, to select one shape (shape ID) to which the naming information obtained from the naming detection unit 51 will be assigned, according to a categorical distribution. For example, if the naming information is "eye", the probability structure is such that horizontal ellipses account for 60%, the shape of a waning crescent moon for 30%, and a circle for 10%.
[0044] Furthermore, when calculating probabilities here, the shape selection unit 52 narrows its calculations to only the shapes held by the dialogue system 20. That is, the shape selection unit 52 normalizes the shapes included in the dialogue so that the sum is "1". In this way, by using the probability structure, which is statistical information, the shape selection unit 52 identifies the shape to be named, making it possible to associate naming information with shapes in accordance with the human intuition of those involved with the data that formed the basis of the statistical information. For example, if there are multiple shape IDs corresponding to the shape of a shape, such as when there are multiple "ellipses", the shape selection unit 52 randomly selects a predetermined shape ID from among the multiple shape IDs with the same probability, according to a uniform distribution.
[0045] Then, the shape selection unit 52 updates the naming information associated with the shape ID managed by the naming information management unit 23 with the naming information obtained from the naming detection unit 51.
[0046] <Target Naming Detection Unit> The target naming detection unit 61 uses the language model M to determine whether the utterance (text) contains target naming information that indicates the goal for completing the arrangement of each shape. If it does, it detects (extracts) the target naming information from the utterance and outputs this target naming information to the target shape arrangement creation unit 62. For example, the target naming detection unit 61 extracts target naming information ("house") from the user's utterance ("Let's build a house").
[0047] <Target Shape Placement Creation Section> The target shape placement creation unit 62 uses the language model M to create a blueprint as a layout for placing each shape by setting the position (placement) coordinates of each shape to be moved so as to approach the completed target indicated by the target naming information, based on the target naming information obtained from the target naming detection unit 61. With such a blueprint, the dialogue system 20 can prevent the placement of shapes by the user from being significantly different from the placement of the dialogue system 20 itself.
[0048] Furthermore, the target shape arrangement creation unit 62 stores and manages the created design drawing in the target information management unit 24, and also outputs it to the moving shape determination unit 63.
[0049] <Moving Shape Determination Unit> The moving shape determination unit 63 compares the current shape arrangement with the shape arrangement of the design drawing obtained from the target shape arrangement creation unit 62, calculates predetermined distances such as Euclidean distances between the coordinates of each shape in the current arrangement and the coordinates of each shape in the design drawing, and determines the one shape that takes the maximum value as the shape to be moved.
[0050] Furthermore, the moving shape determination unit 63 assigns naming information to the determined shape by selecting naming information from the sub-words of the thesaurus based on the target naming information. For example, for "house," the moving shape determination unit 63 selects sub-words of the thesaurus such as "window" or "roof."
[0051] The moving shape determination unit 63 then associates the determined shape (shape ID) with the selected naming information and stores and manages it as naming setting information in the naming setting information management unit 25. At this stage, the naming information managed by the naming information management unit 23 is not updated.
[0052] <Speech content generation unit> The speech content generation unit 71 uses a language model M to generate system speech content that indicates what the dialogue system 20 will say to the user, based on coordinate information read from the coordinate information management unit 21, naming information read from the naming information management unit 23, target information read from the target information management unit 24, and naming setting information read from the naming setting information management unit 25.
[0053] The speech content generation unit 71 then outputs the system speech content as user speech content to the shape movement detection unit 41, the naming detection unit 51, and the target naming detection unit 61, thereby completing one cycle of the above-mentioned processes, and simultaneously transmitting the system speech content data from the communication unit 31 to the user terminal 80. This completes one dialogue as a pair of dialogue content consisting of user speech content and the system speech content in response. For the second and subsequent dialogues, the above-mentioned process is repeated. As a result, the dialogue system 20 gradually adjusts the coordinates of the shape placement in the coordinate information management unit 21 to match the shape placement on the user's side through dialogue with the user. Therefore, the dialogue system 20 can perform the shape placement task together with the user while performing naming. Although a collaborative shape placement task is considered here, the functions described here are also applicable to collaborative work between the user and the system in general, such as work involving naming.
[0054] Furthermore, if the naming information of a figure changes as a result of generating system utterances, the speech content generation unit 71 updates the naming information of the corresponding figure ID in the naming information management unit 23. The speech content generation unit 71 also stores the utterances and system utterances in the speech history management unit 26.
[0055] [Processing of the dialogue system] Next, we will explain the processing of the dialogue system using Figures 4 and 5. Figures 4 and 5 are flowcharts showing the processing of the dialogue system.
[0056] S11: As shown in Figure 1, the communication unit 31 acquires (receives) data of user utterances that have been entered and sent from the user terminal 80 by the user through text input or speaking.
[0057] S12: The shape movement detection unit 41 uses the language model M to obtain predetermined naming information for a predetermined shape and the destination coordinates of this shape from the user utterance (text). The shape movement detection unit 41 also searches the naming information management DB 23 using the predetermined naming information obtained from the language model M as a search key, and identifies the shape by reading the corresponding predetermined shape ID.
[0058] S13: The shape movement unit 42 moves the shape corresponding to the predetermined shape ID to the destination coordinates based on the predetermined shape ID and destination coordinates obtained from the shape movement detection unit 41.
[0059] S14: The shape movement unit 42 updates the coordinate information management unit 21 with the position (placement) coordinates corresponding to the moved predetermined shape ID as the destination coordinates.
[0060] S15: The naming detection unit 51 uses the language model M to determine whether or not naming information is included in the user utterance (text), and if it is included, it detects specific naming information from the user utterance.
[0061] S16: The shape selection unit 52 uses a probability structure managed by the probability structure management unit 22, which represents which shape (shape ID) the naming information can refer to, to select one shape (shape ID) to which the specific naming information obtained from the naming detection unit 51 is to be assigned.
[0062] S17: The shape selection unit 52 updates the naming information associated with the shape ID managed by the naming information management unit 23 with specific naming information obtained from the naming detection unit 51.
[0063] S18: As shown in Figure 5, the target naming detection unit 61 uses the language model M to determine whether the user utterance (text) contains target naming information that indicates the goal for completing the task by placing each figure. If it does, it detects the target naming information from the user utterance.
[0064] S19: The target shape placement creation unit 62 uses the language model M and, based on the target naming information obtained from the target naming detection unit 61, sets the position (placement) coordinates of each shape to be moved so as to approach the completed target indicated by the target naming information, thereby creating a design drawing as a layout for placing each shape.
[0065] S20: The target shape arrangement creation unit 62 stores and manages the created design drawings in the target information management unit 24.
[0066] S21: The moving shape determination unit 63 compares the current shape arrangement (layout) with the shape arrangement (layout) of the design drawing obtained from the target shape arrangement creation unit 62, calculates predetermined distances such as Euclidean distance between the coordinates of each shape in the current state and the coordinates of each shape in the design drawing, and determines the one shape that takes the maximum value as the shape to be moved.
[0067] S22: The moving figure determination unit 63 selects naming information from the sub-words of the thesaurus based on the target naming information for the determined figure, associates the determined figure (figure ID) with the selected naming information, and stores and manages it as naming setting information in the naming setting information management unit 25.
[0068] S23: The speech content generation unit 71 uses the language model M to generate system speech content that indicates what the dialogue system 20 will say to the user (user terminal 80), based on the coordinate information read from the coordinate information management unit 21, the naming information read from the naming information management unit 23, the target information read from the target information management unit 24, and the naming setting information read from the naming setting information management unit 25.
[0069] S24: When the speech content generation unit 71 generates system speech content and the naming information of a figure changes, it updates the naming information of the corresponding figure ID in the naming information management unit 23. The speech content generation unit 71 also stores user speech content and system speech content in the speech history management unit 26.
[0070] S25: The communication unit 31 transmits data of the system utterance to the user terminal 80. As a result, as shown in Figure 1, the user terminal 80 either displays text indicating the system utterance on the display device 1007, or the voice output device 1009 outputs the system utterance as voice.
[0071] S26: The speech content generation unit 71 determines whether the cycle of processing S12-S25 as a system utterance in response to the user utterance has been completed. If it has been completed (YES), processing is temporarily terminated until the communication unit 31 receives data for the next user utterance from the user.
[0072] S27: On the other hand, if the process in S27 is not completed (NO), the speech content generation unit 71 outputs the system speech content it has generated as user speech content to the figure movement detection unit 41, the naming detection unit 51, and the target naming detection unit 61, thereby causing the figure to move in accordance with the system speech content.
[0073] The processing groups S12-S14, S15-S17, and S18-S22 described above may be performed in any order. [Specific examples] Next, a specific example of this embodiment will be described using Figures 6 to 9. Here, an example of actually performing a collaborative figure placement task will be described. In this example, a large-scale language model that has been tuned for instruction will be mainly used as an example of language model M. Here, GPT-4 will be used via API (Application Programming Interface).
[0074] The following process uses a feature called Function Calling in GPT-4. Function Calling is a feature that allows you to pass the definition of a function you have created when calling an API, and it returns the passed function and the arguments required for that function. • Acquisition of "naming information" and "coordinates of the destination" from the speech content in the shape movement detection unit 41. • Acquisition of "naming information" contained in the utterance content by the naming detection unit 51 • Acquisition of "target naming information" from the utterance content in the target naming detection unit 61. • Creation of a "design drawing" from target naming information in the target shape arrangement creation unit 62. • Generation of "system utterance content" from each management information in the utterance content generation unit 71. (1) Detection of shape movement, etc. For example, Figure 6 shows an example of a function calling description in the shape movement detection unit 41. Figure 6 is a diagram showing an example of a function calling description in the shape movement detection unit.
[0075] By using this method, naming information and coordinate information indicating where the figure should be moved can be obtained from the utterance. Of course, it is also possible to obtain the corresponding coordinate information by performing morphological analysis, syntactic analysis, or semantic analysis without using function calling, identifying the word sequence corresponding to the naming information, and further identifying the destination information for the figure (e.g., "right," "down"), and referring to a pre-prepared table. The naming detection unit 51, target naming detection unit 61, and target figure placement creation unit 62 are also implemented using function calling in a similar manner.
[0076] (2) Probability structure For the probabilistic structure, we will use the collaborative figure placement corpus disclosed in Reference 1. This corpus holds naming information and information about the figures (figure IDs) indicated by the naming information for dialogue data from collaborative figure placement tasks. From this, we can obtain P(figure|naming).
[0077] (3) Generation of speech content Two types of prompts are used for generating speech content by the speech content generation unit 71. The first is a prompt used when the dialogue system 20 does not have a target such as a finished product from the beginning of the dialogue (see Figure 7). The second is a prompt used when the dialogue system 20 has a target such as a finished product from the beginning (see Figure 8). When the dialogue system 20 has a target from the beginning of the dialogue, the developer of the dialogue system 20 provides the target to the dialogue system 20, the target shape arrangement creation unit 62 creates a design drawing based on the target from the developer, and the moving shape determination unit 63 performs naming settings. In addition, when naming the target shape determined by the moving shape determination unit 63, the Japanese vocabulary system developed by NTT (Nippon Telegraph and Telephone Corporation) may be used as a thesaurus.
[0078] [Experimental results] Figure 9 shows an example of a collaborative figure placement task performed by having a system with the former prompt and a system with the latter prompt interact. If we consider speaker A to correspond to the dialogue system 20 and speaker B to correspond to the user, it can be confirmed that similar figure placements can be achieved through naming information. Therefore, this embodiment shows that collaborative work between the user and the dialogue system 20 is possible through naming information.
[0079] [Main effects of this embodiment] As described above, according to this embodiment, if the dialogue system 20 can name the shapes, then efficient communication can be achieved between the user (human) and the dialogue system 20, just as it is between humans.
[0080] Furthermore, when there are multiple objects and the user wants to share the selected object with the dialogue system 20, this action can be performed efficiently. For example, in tasks where multiple objectives are given, such as furniture placement, navigation, or finding items, a dialogue system can be realized that allows for the efficient sharing of selected objects.
[0081] 〔supplement〕 The present invention is not limited to the embodiments described above, and may also have the following configurations or processes (operations).
[0082] (1) The dialogue system 20 can be implemented using a computer and a program, but this program can also be recorded on a (non-temporary) recording medium or provided via a communication network such as the Internet.
[0083] (2) The processor 1001 may be single or multiple. [Explanation of Symbols]
[0084] 10 Communication Systems 20 Dialogue Systems 21 Coordinate information management department 22 Probability Structure Management Department 23 Naming Information Management Department 24 Target Information Management Department 25. Naming Settings Information Management Department 26. Speech History Management Department 31 Communications Department 41. Shape movement detection unit 42 Shape Movement Section 51 Naming detection unit 52 Shape Selection Section 61 Target Naming Detection Unit 62 Target Shape Placement Creation Section 63 Moving Figure Determination Unit 71. Speech content generation unit
Claims
1. An interactive system that generates geometric object placement coordinates by arranging geometric objects while interacting with a user operating a user terminal, A communication unit that receives user utterances sent from the user terminal by the user, A graphic movement detection unit uses a language model to obtain predetermined naming information for a predetermined graphic object and the destination coordinates of the ID of the predetermined graphic object from the user utterance content, and reads the ID of the graphic object corresponding to the predetermined naming information from a naming information management unit which manages the IDs and naming information of each graphic object. A coordinate information management unit manages the ID and position coordinates of each geometric object, and a geometric object movement unit updates the position coordinates corresponding to the ID of a predetermined geometric object to the destination coordinates. A dialogue system having
2. The dialogue system according to claim 1, A naming detection unit that uses the language model to determine whether or not naming information is included in the user utterance, and if it is included, detects specific naming information from the user utterance; A graphic selection unit selects the ID of a specific graphic object to which the naming information can represent using a probability structure that indicates which graphic object ID the naming information can represent, and updates the naming information associated with the ID of the specific graphic object managed by the naming information management unit to the specific naming information. A dialogue system having
3. The dialogue system according to claim 2, A target naming detection unit uses the language model to determine whether the user utterance contains target naming information indicating the goal for arranging each graphic object to complete the task, and if it does, detects the target naming information from the user utterance. A target shape placement creation unit creates a layout for a graphic object to be moved by setting the position coordinates of the graphic object to be moved so that it approaches the target indicated by the target naming information, using the language model described above. A moving shape determination unit compares the current layout with the layout newly created by the target shape placement creation unit, calculates the distance between the position coordinates of each shape object in the current layout and the position coordinates of each shape object in the newly created layout, and determines the one shape object that takes the maximum distance as the shape object to be moved. A dialogue system having
4. The dialogue system according to claim 3, A dialogue system having a speech content generation unit that generates system speech content indicating what the dialogue system will speak to the user terminal, based on the position coordinates read from the coordinate information management unit, the naming information read from the naming information management unit, and the naming setting information determined by the moving figure determination unit, using the language model.
5. The dialogue system according to claim 4, wherein the speech content generation unit outputs the system speech content as the user speech content to the figure movement detection unit, the naming detection unit, and the target naming detection unit.
6. A dialogue method performed by a dialogue system that generates geometric object placement coordinates by interacting with a user operating a user terminal, Communication processing to receive user utterance content sent from the user terminal by the user, Using a language model, determine whether or not naming information is included in the user utterance, and if so, perform a naming detection process to detect specific naming information from the user utterance. A shape selection process that uses a probability structure representing which shape object ID the naming information can represent to select the ID of a specific shape object to which the specific naming information is to be assigned, and updates the naming information associated with the ID of the specific shape object managed in the naming information management unit, which manages the IDs and naming information of each shape object, to the specific naming information. A method of interaction to perform this action.
7. A dialogue method performed by a dialogue system that generates geometric object placement coordinates by interacting with a user operating a user terminal, Communication processing to receive user utterance content sent from the user terminal by the user, Using a language model, determine whether the user utterance contains goal naming information indicating the goal for completing the task by placing each graphic object. If it does, perform a goal naming detection process to detect the goal naming information from the user utterance. Using the language model, a target shape placement creation process creates a layout for the target shape object by setting the position coordinates of the shape object to be moved so that it approaches the target indicated by the target naming information, based on the target naming information. A moving shape determination process that compares the current layout with the layout newly created by the target shape placement creation process, calculates the distance between the position coordinates of each shape object in the current layout and the position coordinates of each shape object in the newly created layout, and determines the shape object that takes the maximum distance as the shape object to be moved. A method of interaction to perform this action.
8. A program that causes a computer to perform the method described in claim 6 or 7.