Speech dialogue system
The voice dialogue system addresses user discomfort by seamlessly integrating machine learning and human responses with common voice parameters and avatars, improving user experience and reducing operator workload.
Patent Information
- Application Number
- JP2024107483
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2026-01-16
AI Technical Summary
Existing voice dialogue systems that combine responses from machine learning models and human operators can cause user discomfort due to subtle differences in response styles, making it difficult for users to distinguish between the two.
A voice dialogue system that includes a voice dialogue unit and a control unit, capable of executing either a first dialogue process using a machine learning model or a second dialogue process using an operator, with common voice parameters and avatars to mask the difference, and adjusts dialogue based on user and operator states and preferences.
The system reduces user discomfort by making it difficult to distinguish between machine learning model and operator responses, enhancing the user experience and reducing operator burden.
Smart Images

Figure 2026007538000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a voice dialogue system. [Background technology]
[0002] There is known a system in which users of facilities such as commercial facilities have a voice conversation with an operator outside the facility (for example, Patent Document 1). By introducing such a system, the operator can remotely respond to user requests.
[0003] On the other hand, chatbots that respond to user requests in natural language using machine learning models such as large-scale language models are also known. In such systems, chatbots can also be used instead of operators. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-149630 Summary of the Invention [Problem to be solved by the invention]
[0005] However, there are subtle differences between responses based on machine learning models and responses by operators. Therefore, in a system that responds to user requests, using both responses based on machine learning models and responses by operators may cause discomfort to the user. [Means for solving the problem]
[0006] (1) A voice dialogue system that solves the problem is a voice dialogue system that includes: a voice dialogue unit that outputs to a user a dialogue voice generated by a voice synthesis engine that generates a dialogue voice based on specified voice parameters from sentence information including a response text; and a control unit that executes a first dialogue process that causes the voice dialogue unit to output a first dialogue voice generated from first sentence information, and a second dialogue process that causes the voice dialogue unit to output a second dialogue voice generated from second sentence information, wherein the first sentence information includes a response text regarding a response to the user that is generated by a machine learning model that generates sentence information by natural language processing, and the second sentence information includes a response text regarding a response by the operator to the user that is input to an operator terminal operated by the operator, and the control unit executes either the first dialogue process or the second dialogue process according to predetermined selection conditions.
[0007] According to the above configuration, dialogue speech adjusted based on common voice parameters is output in the first dialogue process and the second dialogue process, making it difficult to distinguish between the first dialogue speech and the second dialogue speech. Therefore, when a user engages in a voice dialogue with the voice dialogue system, it is difficult for the user to tell whether the dialogue is with a machine learning model or with an operator. This makes it possible for the voice dialogue system to reduce the sense of discomfort felt by the user during the voice dialogue between the voice dialogue system and the user.
[0008] (2) In the voice dialogue system described in (1) above, in the first dialogue processing, the control unit causes the machine learning model to generate original text regarding a response to the user's request, causes the machine learning model to form response text by converting the original text based on sentence parameters regarding the structure of the sentence, forms the first sentence information including the response text based on the original text, and outputs the first sentence information to the voice dialogue unit.
[0009] The original text generated by the machine learning model may be unnatural compared to the language used by the operator. According to the above configuration, the original text generated by the machine learning model is converted into response text based on the sentence parameters. Then, the first sentence information is formed based on the response text. Therefore, the response text included in the first sentence information is not the sentence itself generated by the machine learning model. Therefore, according to the above first dialogue processing, it is possible to suppress the user's sense of discomfort.
[0010] (3) In the voice dialogue system described in (1) or (2) above, the operator's voice is input to the operator terminal, and the control unit, in the second dialogue processing, forms a response text by converting the operator's voice into text, forms the second sentence information including the response text based on the operator's voice, and outputs the second sentence information to the voice dialogue unit.
[0011] With this configuration, the operator's voice is converted into text once, and therefore the operator's characteristics such as voice quality, volume, and emotion are removed from the second dialogue voice. By removing the operator's unique characteristics from the text on which the second dialogue voice is based, it becomes difficult for the user to tell whether the voice is based on a machine learning model or the operator.
[0012] (4) In the voice dialogue system according to any one of (1) to (3) above, the control unit executes either the first dialogue process or the second dialogue process according to operator information relating to the state of the operator.
[0013] According to this configuration, either the first interaction process or the second interaction process can be executed depending on the operator information.
[0014] (5) In the voice dialogue system described in any one of (1) to (4) above, the control unit executes the first dialogue processing when the voice dialogue is started in response to at least one of an input action by the user and the state of the user, and executes the second dialogue processing when the voice dialogue is started by operation of the operator terminal.
[0015] According to this configuration, when a voice dialogue is started by at least one of a user's input action and a user's state, the first dialogue process is executed, so that the machine learning model can automatically respond to the user's request. On the other hand, when a voice dialogue is started by an operation of the operator terminal, the operator can respond to the user's request.
[0016] (6) In the voice dialogue system described in any one of (1) to (5) above, the control unit executes either the first dialogue processing or the second dialogue processing depending on user information relating to attributes of the user.
[0017] According to this configuration, either the first interaction process or the second interaction process can be executed depending on the user information.
[0018] (7) In the voice dialogue system described in any one of (1) to (6) above, the voice dialogue unit has a display that displays an avatar to the user, and the avatar common to the first dialogue processing and the second dialogue processing is displayed on the display.
[0019] According to this configuration, a common avatar is displayed on the display in both the first and second dialogue processes, which makes it difficult for the user to tell whether the voice dialogue is with a machine learning model or with an operator when the user engages in a voice dialogue with the voice dialogue system.
[0020] (8) In the voice dialogue system described in any one of (1) to (7) above, the control unit generates a re-learning dataset for the machine learning model based on the second sentence information, and updates the machine learning model by re-learning the re-learning dataset to the machine learning model.
[0021] According to this configuration, the machine learning model is updated with the re-learning dataset generated based on the second sentence information, so that the original text generated by the machine learning model becomes closer to the operator's answer, thereby reducing the sense of discomfort felt by the user in the dialogue between the voice dialogue system and the user. [Effects of the Invention]
[0022] According to the voice dialogue system, it is possible to reduce the sense of discomfort felt by the user in the voice dialogue between the voice dialogue system and the user. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a schematic configuration diagram illustrating an example of the configuration of a voice dialogue system according to an embodiment. [Figure 2] 1 is a block diagram showing an example of the configuration of a voice dialogue system according to an embodiment; [Figure 3] FIG. 10 is a block diagram showing the flow of a first dialogue process. [Figure 4] FIG. 10 is a block diagram showing the flow of a second dialogue process. [Figure 5] 4 is a flowchart illustrating a process executed by a control unit according to the embodiment. [Figure 6] 10 is a flowchart showing a process executed by a control unit in a voice dialogue system according to a first modified example. [Figure 7] FIG. 10 is a block diagram showing the flow of a second dialogue process in a voice dialogue system according to a second modified example. [Figure 8] FIG. 11 is a block diagram showing the flow of a second dialogue process in a voice dialogue system according to a third modified example. DETAILED DESCRIPTION OF THE INVENTION
[0024] <Embodiment> A voice dialogue system 10 according to this embodiment will be described with reference to FIGS. The voice dialogue system 10 shown in Fig. 1 is a system for voice dialogue with users who use a facility. The facility is a facility that provides some kind of service to users of the facility, such as a commercial facility, public facility, educational facility, medical facility, or nursing home. Examples of facilities include retail stores, accommodation facilities, government buildings, schools, hospitals, nursing homes, restaurants, financial institutions, tourist destinations, theme parks, commercial buildings, airports, train stations, and bus stops.
[0025] As shown in FIGS. 1 and 2, a voice dialogue system 10 engages in a voice dialogue with a user. In the voice dialogue system 10, a user request is responded to either automatically by a machine learning model 110 or by an operator. The voice dialogue system 10 is connected to the machine learning model 110 and the operator via a network. Specifically, a voice dialogue unit 20 of the voice dialogue system 10 installed in a facility is connected via the network to an external server 100 including the machine learning model 110 and an operator terminal 200 operated by the operator. To make it easier for the voice dialogue system 10 to interact with the user, an avatar is set in the voice dialogue system 10. When a user engages in a voice dialogue using the voice dialogue system 10, the user feels as if they are interacting with an avatar, rather than with either the machine learning model 110 or the operator.
[0026] The voice dialogue system 10 allows an operator to respond to a user's request remotely, thereby contributing to a variety of working styles for the operator. Furthermore, because the operator has a voice dialogue with the user via an avatar, the burden on the operator that would otherwise be required to directly interact with the user is reduced. The voice dialogue system 10 can also automatically respond to a user's request using the machine learning model 110. By automatically responding using the machine learning model 110, the operator does not have to respond directly to the user's request, thereby reducing the burden on the operator.
[0027] On the other hand, there are subtle differences between the responses provided by the machine learning model 110 and the responses provided by the operator. Therefore, using both the responses provided by the machine learning model 110 and the responses provided by the operator may cause the user to feel uncomfortable. The voice dialogue system 10 can reduce this feeling of discomfort. The voice dialogue system 10 will be described below.
[0028] <Spoken dialogue system> As shown in FIG. 2, the voice dialogue system 10 includes a voice dialogue unit 20 and a control unit 30.
[0029] The voice dialogue unit 20 is a part of the voice dialogue system 10 that is arranged on the front-end side 11. The voice dialogue unit 20 is installed, for example, in a facility. The voice dialogue unit 20 is configured, for example, as a computer terminal. Examples of computer terminals include a tablet, a personal computer, and a smart speaker.
[0030] The control unit 30 is a part of the voice dialogue system 10 that is arranged on the back-end side 12. The control unit 30 is included in, for example, a computer that operates according to a predetermined program. The computer including the control unit 30 is mounted, for example, on a computer terminal that is the voice dialogue unit 20. The computer including the control unit 30 may be provided separately from the computer terminal that is the voice dialogue unit 20. When the computer including the control unit 30 is provided separately from the computer terminal that is the voice dialogue unit 20, the computer including the control unit 30 may be configured as a server. The server is configured by one or more server devices. The server is connected to the computer terminal that is the voice dialogue unit 20 via a network, for example. The server may be a cloud server.
[0031] <Voice dialogue section> The voice dialogue unit 20 shown in FIG. 2 has a microphone 21, a camera 22, and an input device 23. The microphone 21, the camera 22, and the input device 23 acquire information about the inside of the facility and the user. The microphone 21 acquires sounds inside the facility. The microphone 21 acquires the user's voice during a voice dialogue with the user. The camera 22 captures images inside the facility. The camera 22 captures images of the user during a voice dialogue with the user. The images captured by the camera 22 are transmitted to the control unit 30. Information can be input into the input device 23 through user operation. The input device 23 may be a keyboard, a mechanical switch, or a touch panel provided on the display 25.
[0032] The voice interaction unit 20 has a speaker 24. The speaker 24 outputs voice into the facility. The speaker 24 outputs voice to the user during a voice interaction with the user. The voice interaction unit 20 has a display 25. The display 25 displays an avatar to the user. The avatar displayed on the display 25 changes its gestures, facial expressions, and other elements in accordance with the voice output from the speaker 24, for example. The avatar may be an animated character or a realistic character like a live-action character. Using a realistic image of a person as the avatar makes the user feel as if they are being noticed by the operator. The display 25 may display text related to the voice output from the speaker 24. The display 25 may also display the waiting time for a response to the user's request.
[0033] The voice dialogue unit 20 has a voice synthesis engine 26. The voice dialogue unit 20 outputs a dialogue voice generated by the voice synthesis engine 26 to the user from a speaker 24. The voice synthesis engine 26 generates a dialogue voice based on specified voice parameters from text information including a response text. In addition to the text, the text information may include information related to the text, such as the time the text was generated, information about the avatar's movements, and a URL related to the text. Information related to the text of the text information may be displayed on a display 25.
[0034] The voice parameters are set in advance according to the avatar character. The voice synthesis engine 26 generates dialogue voice by reading out the text included in the text information based on the specified voice parameters. The voice parameters include, for example, elements such as volume, pitch, speed, range, and voice quality of the voice reading the text.
[0035] The voice dialogue unit 20 has a sentence generation unit 27. The sentence generation unit 27 generates text (hereinafter referred to as user text) from the user's voice acquired by the microphone 21. The sentence generation unit 27 has, for example, a voice recognition engine. The user's voice acquired by the microphone 21 is converted into user text by the sentence generation unit 27. The voice dialogue unit 20 forms sentence information including user text generated from the user's voice by voice-to-text conversion processing. The voice dialogue unit 20 transmits the sentence information to the control unit 30. The control unit 30 may have the sentence generation unit 27.
[0036] <Control unit> 2 controls the operation of the voice dialogue system 10. The control unit 30 includes, for example, a processor. The processor includes, for example, at least one of a central processing unit and a microprocessor.
[0037] The voice dialogue system 10 includes, for example, a storage unit 31. The storage unit 31 stores programs related to the control executed by the control unit 30. The storage unit 31 includes, for example, a non-volatile, rewritable semiconductor memory. The control unit 30 executes the programs stored in the storage unit 31 to perform various processes.
[0038] <Machine learning model> The voice dialogue system 10 shown in Fig. 2 is connected to an external server 100, for example, via a network. Examples of networks include the Internet, a local network, a telephone line network, a dedicated line network such as RS-485, and a composite network formed by connecting these. The same applies to the following description of a network. The external server 100 is composed of one or more server devices. The external server 100 may be located in the same facility as the voice dialogue system 10, or in a different facility from the voice dialogue system 10.
[0039] The external server 100 includes a machine learning model 110. The machine learning model 110 is, for example, a large-scale language model capable of natural language processing. As the machine learning model 110, a generation AI such as ChatGPT, Bard (registered trademark), or Claude can be used.
[0040] <Operator terminal> The voice dialogue system 10 shown in Fig. 2 is connected to an operator terminal 200 operated by an operator via a network, for example. The operator's voice is input to the operator terminal 200. The operator terminal 200 may be located in the same facility as the voice dialogue system 10, or may be located in a facility different from the voice dialogue system 10. The operator terminal 200 is, for example, a computer terminal such as a personal computer, a tablet, a smartphone, or a telephone. The operator terminal 200 has a communication device 210, a display unit 220, and an input unit 230.
[0041] The communication device 210 is a device for an operator to have a voice dialogue with a user. The communication device 210 has a microphone that captures the voice of the operator. The voice of the operator captured by the microphone of the communication device 210 is transmitted to the control unit 30 of the voice dialogue system 10. The communication device 210 may have a speaker that outputs the voice of the user. The communication device 210 may be configured as a headset.
[0042] The display unit 220 has, for example, a display. For example, the display unit 220 displays an image captured by the camera 22 of the voice dialogue unit 20 and text related to the user's voice.
[0043] The input unit 230 is operated by an operator and includes an input device such as a keyboard and a touch panel provided on the display of the display unit 220.
[0044] <Voice dialogue> The control unit 30 shown in Fig. 2 executes a first dialogue process and a second dialogue process. By executing either the first dialogue process or the second dialogue process, a voice dialogue between the voice dialogue system 10 and a user is executed. The first dialogue process is a process of conducting a voice dialogue with a user through an automatic response by the machine learning model 110. The second dialogue process is a process of conducting a voice dialogue with a user through a response by an operator.
[0045] The voice dialogue system 10 is configured to make it difficult for the user to determine whether the first dialogue processing or the second dialogue processing is being executed during a voice dialogue with the user. Voice parameters common to the first dialogue processing and the second dialogue processing are set as voice parameters used to generate dialogue voice in the voice dialogue system 10. An avatar common to the first dialogue processing and the second dialogue processing is displayed on the display 25 of the voice dialogue unit 20. This makes it difficult for the user to determine whether the voice dialogue with the voice dialogue system 10 is an automatic response using the machine learning model 110 or a response by an operator.
[0046] <First dialogue process> The first dialogue processing is a process of causing the voice dialogue unit 20 to output a first dialogue voice generated from first sentence information. The first sentence information includes a response text regarding a response to the user. The response text included in the first sentence information is generated by a machine learning model 110 that generates sentence information by natural language processing. The response text regarding a response to the user generated by the machine learning model 110 is generated by inputting user text regarding the user's voice to the machine learning model 110 as a prompt.
[0047] Specifically, in the first dialogue processing, the control unit 30 causes the machine learning model 110 to generate original text regarding a response to a user request. The control unit 30 causes the machine learning model 110 to form response text by converting the original text based on sentence parameters related to the structure of the sentence. The control unit 30 forms first sentence information including the response text based on the original text. Then, the control unit 30 outputs the first sentence information to the voice dialogue unit 20.
[0048] The original text generated by the machine learning model 110 may be unnatural because it is long, has a large number of words, is written in an explanatory tone, etc. The sentence parameters are set from the original text generated by the machine learning model 110 so as to suppress unnaturalness caused by the generation by the machine learning model 110.
[0049] The text parameters include, for example, elements such as the number of characters in the text, the number of words, tone of voice, endings, first person, etc. The text parameters may be set according to the character of the avatar. The text parameters may be set to elements appropriate for a conversation between a service representative who serves users at a facility and the facility user. As an example of appropriate elements, if the facility is a retail store, the text parameters may include elements such as tone of voice, endings, first person, etc. that do not make the customer (shopper) feel rude when the service representative (store clerk) speaks to the customer (shopper) of the retail store.
[0050] The control unit 30 converts the original text into a response text, for example, by inputting a prompt including the original text and sentence parameters generated by the machine learning model 110 to the machine learning model 110. The process of converting the original text based on the sentence parameters may be performed by a language processing device that performs natural language processing, different from the machine learning model 110. A generation AI different from the machine learning model 110 may be used as the language processing device that performs natural language processing.
[0051] The flow of the first interaction process will be described with reference to FIG. First, in step S11, the sentence generation unit 27 generates user text from the user's voice acquired by the microphone 21 of the voice dialogue unit 20. The control unit 30 inputs sentence information including the user text to the machine learning model 110.
[0052] In step S12, the machine learning model 110 generates original text as a response to the user request based on the user text. In step S13, the machine learning model 110 converts the generated original text based on the sentence parameters. The control unit 30 forms first sentence information so as to include response text formed based on the sentence parameters. The control unit 30 outputs the first sentence information to the voice dialogue unit 20. Steps S12 and S13 may be performed simultaneously by one prompt.
[0053] In step S14, the voice synthesis engine 26 generates a first dialogue voice from the first sentence information. Specifically, the voice synthesis engine 26 generates the first dialogue voice from the response text included in the first sentence information. In step S15, the voice dialogue unit 20 outputs the first dialogue voice from the speaker 24.
[0054] <Second dialogue processing> The second dialogue processing is a process of causing the voice dialogue unit 20 to output a voice for the second dialogue generated from the second sentence information. The second sentence information includes a response text regarding the response by the operator to the user, which is input to the operator terminal 200. The response text regarding the response by the operator to the user is a text converted from the voice of the operator input to the operator terminal 200.
[0055] Specifically, in the second dialogue processing, the control unit 30 forms a response text included in the second sentence information by converting the operator's voice into text. The control unit 30 has a voice recognition engine for converting the operator's voice into text. The process of converting the operator's voice into text may be performed in the operator terminal 200. In the second dialogue processing, the control unit 30 forms the response text by converting the operator's voice into text. The control unit 30 forms second sentence information including the response text based on the operator's voice. Then, the control unit 30 outputs the second sentence information to the voice dialogue unit 20.
[0056] The flow of the second interaction process will be described with reference to FIG. First, in step S21, the sentence generation unit 27 generates user text from the user's voice acquired by the microphone 21 of the voice dialogue unit 20. The control unit 30 outputs sentence information including user text related to the user's voice to the operator terminal 200.
[0057] In step S22, the operator terminal 200 acquires the operator's response. Specifically, user text related to the user's voice is displayed on the display unit 220 of the operator terminal 200. When the operator who reads the user text responds to the user's request by voice, the operator terminal 200 acquires the operator's response by voice from the communication device 210. The operator terminal 200 outputs voice data including the voice related to the operator's response to the control unit 30 of the voice dialogue system 10.
[0058] In step S23, the control unit 30 converts the operator's response into text. Then, the control unit 30 forms a response text through this conversion. The control unit 30 generates second sentence information including the response text based on the operator's response. The control unit 30 outputs the second sentence information to the voice dialogue unit 20.
[0059] In step S24, the voice synthesis engine 26 generates a second dialogue voice from the second sentence information. Specifically, the voice synthesis engine 26 generates the second dialogue voice from the response text included in the second sentence information. In step S25, the voice dialogue unit 20 outputs the second dialogue voice from the speaker 24.
[0060] <Executing voice dialogue> The control unit 30 starts a voice dialogue with the user in response to at least one of an input action by the user and the user's state. The control unit 30 acquires at least one of an input action by the user and the user's state from the voice acquired by the microphone 21 of the voice dialogue unit 20, the video captured by the camera 22, the input to the input device 23, etc.
[0061] The input action by the user is, for example, at least one of the following: a voice input by the user into the microphone 21, the user appearing in the camera 22, and the user operating the input device 23. The user input action is executed when the user wishes to start a voice dialogue or when the user answers a question from the voice dialogue system 10 during a voice dialogue. The control unit 30 starts the voice dialogue when at least one of the following is performed: a voice input by the user into the microphone 21, the user appearing in the camera 22, and the user operating the input device 23.
[0062] The user's state includes states in which the user is assumed to be in distress, such as when the user's voice contains specific words such as "I'm in trouble" or "I don't understand," when the user has passed the same place multiple times and is lost, or when the user is unable to move due to poor health. The control unit 30 starts a voice dialogue when it determines that the user is in any of the above states based on the voice acquired by the microphone 21, the video captured by the camera 22, the input to the input device 23, etc. The determination of the user's state may be made by the control unit 30 or an operator. In addition, the voice dialogue may be started when the user enters a facility, behaves suspiciously, or is drunk, etc., and the user needs to be addressed.
[0063] The control unit 30 also starts a voice dialogue with the user when the operator terminal 200 is operated. The operation of the operator terminal 200 is, for example, an operation by the operator, such as voice input to the communication device 210 of the operator terminal 200 or operation of the input unit 230. The operator starts a voice dialogue when, for example, the user's state indicates that the user requires a voice dialogue.
[0064] The control unit 30 executes either the first dialogue processing or the second dialogue processing in accordance with a predetermined selection condition. Examples of the selection condition include first to third selection conditions. In the voice dialogue system 10, any one of the first to third selection conditions is set as the selection condition for selecting either the first dialogue processing or the second dialogue processing. At least two of the first to third selection conditions may be set as the selection condition, along with their respective priorities.
[0065] <First selection condition> The first selection condition relates to the status of the operator. An example of the first selection condition is as follows: When the operator is unable to handle voice dialogue, the first dialogue processing is selected, and when the operator is able to handle voice dialogue, the second dialogue processing is selected.
[0066] The control unit 30 executes either the first dialogue process or the second dialogue process according to operator information relating to the status of the operator. The control unit 30 executes the first dialogue process when the operator is unable to respond to voice dialogue. The control unit 30 executes the second dialogue process when the operator is able to respond to voice dialogue. The operator information includes at least one of the operation status of the operator terminal 200 and operator absence information.
[0067] The operation state of the operator terminal 200 is, for example, the operation state of the input unit 230. The input unit 230 is provided with a switch that is operated by the operator. The switch of the input unit 230 is operated when the operator is unable to handle the voice dialogue. When the switch of the input unit 230 is operated, the operator terminal 200 transmits a signal to the voice dialogue system 10 indicating that the operator is unable to handle the voice dialogue. When there is no input to the input unit 230 for a predetermined period of time, the operator terminal 200 may transmit a signal to the voice dialogue system 10 indicating that the operator is unable to handle the voice dialogue.
[0068] The information about the absence of an operator includes schedule information indicating the planned date, time, etc. when the operator is available to respond to a voice interaction. The schedule information is stored, for example, in the storage unit 31. The control unit 30 determines whether the operator is available to respond to a voice interaction based on the schedule information.
[0069] The operator terminal 200 may be equipped with a camera that captures an image of the operator. The operator terminal 200 determines whether the operator is available for voice interaction from the image captured by the camera that captures the operator. The operator terminal 200 determines by image recognition that the operator is not captured in the image captured by the camera. When the operator is not captured in the image captured by the camera, the operator terminal 200 transmits a signal to the voice interaction system 10 indicating that the operator is not available for voice interaction. The determination of whether the operator is not captured in the image captured by the camera may be performed by the control unit 30.
[0070] <Second selection condition> The second selection condition is a condition related to the start of a voice dialogue. An example of the second selection condition is as follows: The second selection condition is to execute the first dialogue process when the voice dialogue is started in response to at least one of an input action by the user and a state of the user, and to execute the second dialogue process when the voice dialogue is started by an operation of the operator terminal 200.
[0071] The control unit 30 executes a first dialogue process when a voice dialogue is started in response to at least one of a user's input operation and the user's state. The control unit 30 executes a second dialogue process when a voice dialogue is started by an operation on the operator terminal 200. The control unit 30 executes the first dialogue process when a voice dialogue is started in response to information acquired by the microphone 21, camera 22, and input device 23 of the voice dialogue unit 20. On the other hand, the control unit 30 executes a second dialogue process when the operator determines that a voice dialogue is necessary and starts the voice dialogue.
[0072] <Third selection condition> The third selection condition is a condition related to the user's state. An example of the third selection condition is as follows: The third selection condition is to execute the first dialogue process if the first dialogue process is suitable for the user's attributes, and to execute the second dialogue process if the second dialogue process is suitable for the user's attributes.
[0073] The control unit 30 executes either the first dialogue process or the second dialogue process according to user information relating to the attributes of the user. The user information includes, for example, at least one of the user's gender and age. The control unit 30 acquires the user information from at least one of the voice acquired by the microphone 21 and the video captured by the camera 22. The control unit 30 acquires the user information from the voice acquired by the microphone 21 using a voice recognition technology. The control unit 30 acquires the user information from the video captured by the camera 22 using an image recognition technology. An operator may determine the user information based on at least one of the voice acquired by the microphone 21 and the video captured by the camera 22. The user information determined by the operator is transmitted to the control unit 30.
[0074] For example, the control unit 30 executes a first dialogue process when the user's age is younger than a predetermined age, and executes a second dialogue process when the user's age is equal to or greater than the predetermined age. The predetermined age is set to, for example, an age at which a user is likely to feel uncomfortable with a voice dialogue using the machine learning model 110.
[0075] The control unit 30 may execute either the first dialogue process or the second dialogue process based on information about whether the user prefers the first dialogue process or the second dialogue process for each age and gender of the user. Such information can be obtained, for example, by analyzing past voice dialogues.
[0076] The control unit 30 may execute either the first interaction process or the second interaction process based on information indicating whether a specific user prefers the first interaction process or the second interaction process. The control unit 30 identifies the user from the user information.
[0077] The user information may include the user's request. If the user requests the first dialogue processing, the control unit 30 executes the first dialogue processing, and if the user requests the second dialogue processing, the control unit 30 executes the second dialogue processing. For example, at the start of a voice dialogue, the control unit 30 confirms with the user whether the user requests the first dialogue processing or the second dialogue processing.
[0078] <Processing in a spoken dialogue system> The process in which the control unit 30 executes either the first dialogue process or the second dialogue process will be described with reference to Fig. 5. After the process in Fig. 5 ends, the control unit 30 executes the process in Fig. 5 again after a predetermined period of time has elapsed.
[0079] In step S31, the control unit 30 determines whether or not to start a voice dialogue. The control unit 30 determines whether or not to start a voice dialogue based on the input operation by the user, the user's state, and the operation of the operator terminal 200. If the voice dialogue is to be started, the control unit 30 proceeds to step S32. If the voice dialogue is not to be started, the control unit 30 ends the processing of FIG. 5.
[0080] In step S32, the control unit 30 determines whether to execute the first dialogue process or the second dialogue process according to a predetermined selection condition. If execution of the first dialogue process is selected, the control unit 30 proceeds to step S33. If execution of the first dialogue process is not selected, the control unit 30 proceeds to step S35. If execution of the first dialogue process is not selected, execution of the second dialogue process is selected.
[0081] In step S33, the control unit 30 executes the first dialogue process, and then the process proceeds to step S34.
[0082] In step S34, the control unit 30 determines whether the voice dialogue has ended. For example, the control unit 30 determines that the voice dialogue has ended if there is no input action by the user for a predetermined period of time after the voice dialogue unit 20 outputs a voice in step S33. If the voice dialogue has ended, the control unit 30 ends the processing of Fig. 5. If the voice dialogue has not ended, the control unit 30 repeats the processing of step S33.
[0083] In step S35, the control unit 30 executes the second dialogue process, and then the process proceeds to step S36.
[0084] In step S36, the control unit 30 determines whether the voice dialogue has ended. For example, the control unit 30 determines that the voice dialogue has ended when there is no input action by the user for a predetermined period of time after outputting a voice from the voice dialogue unit 20 in step S35. The control unit 30 may determine that the voice dialogue has ended when, for example, a signal regarding the end of the voice dialogue is received from the operator terminal 200. For example, when the voice dialogue ends, the operator causes the operator terminal 200 to send a signal regarding the end of the voice dialogue to the control unit 30. If the voice dialogue has ended, the control unit 30 ends the second dialogue processing and terminates the processing of FIG. 5. If the voice dialogue has not ended, the control unit 30 repeats the processing of step S35.
[0085] <Operation of this embodiment> In the voice dialogue system 10, voice is generated using a common voice synthesis engine 26 in both the first dialogue process and the second dialogue process. This makes it difficult for the user to tell whether the other party in the voice dialogue is the machine learning model 110 or an operator. Also, some users may prefer to have everyday conversations with an operator. In such cases, the machine learning model 110 can respond to the user.
[0086] Incidentally, it is known that people behave better when they are receiving attention, such as by observing good manners and not committing crimes, than when they are not. This is called the Hawthorne effect. From the perspective of improving comfort and maintaining public order, it is desirable to achieve the Hawthorne effect in facilities such as commercial facilities. If the voice dialogue is only an automatic response by the machine learning model 110, the user may perceive that the operator is not paying attention to them, and the Hawthorne effect may be reduced.
[0087] In the voice dialogue system 10, it is difficult for the user to distinguish whether the dialogue with the avatar is a dialogue with the machine learning model 110 or a dialogue with an operator. For this reason, the Hawthorne effect becomes stronger when the user recognizes that the operator is paying attention to them.
[0088] <Effects of this embodiment> The effects of this embodiment will be described. (1) The voice dialogue system 10 includes a voice dialogue unit 20 and a control unit 30. The voice dialogue unit 20 outputs to the user a dialogue speech generated by a speech synthesis engine 26 that generates a dialogue speech based on specified voice parameters from text information including a response text. The control unit 30 executes a first dialogue process that causes the voice dialogue unit 20 to output a first dialogue speech generated from first text information, and a second dialogue process that causes the voice dialogue unit 20 to output a second dialogue speech generated from second text information. The first text information includes a response text regarding a response to the user, generated using a machine learning model 110 that generates text information by natural language processing. The second text information includes a response text regarding a response by the operator to the user, which is input to an operator terminal 200 operated by the operator. The control unit 30 executes either the first dialogue process or the second dialogue process according to a predetermined selection condition.
[0089] According to the above configuration, dialogue speech adjusted based on common voice parameters is output in the first dialogue process and the second dialogue process, making it difficult to distinguish between the first dialogue speech and the second dialogue speech. Therefore, when a user engages in a voice dialogue with the voice dialogue system 10, it is difficult for the user to tell whether the dialogue is with the machine learning model 110 or with an operator. As a result, the voice dialogue system 10 can reduce the sense of discomfort felt by the user in the voice dialogue between the voice dialogue system 10 and the user.
[0090] (2) In the first dialogue processing, the control unit 30 causes the machine learning model 110 to generate original text regarding a response to a user request. The control unit 30 causes the machine learning model 110 to form response text by converting the original text based on sentence parameters related to the structure of the sentence. The control unit 30 forms first sentence information including the response text based on the original text. The control unit 30 outputs the first sentence information to the voice dialogue unit 20.
[0091] The original text generated by the machine learning model 110 may be unnatural compared to the language used by the operator. According to the above configuration, the original text generated by the machine learning model 110 is converted into response text based on the sentence parameters. Then, first sentence information is formed based on the response text. Therefore, the text included in the first sentence information is not the sentence itself generated by the machine learning model 110. Therefore, according to the above first dialogue processing, it is possible to suppress the user's sense of discomfort.
[0092] (3) In the second dialogue process, the control unit 30 generates a response text by converting the operator's voice into text. The control unit 30 generates second sentence information including the response text based on the operator's voice. In the second dialogue process, the control unit 30 outputs the second sentence information to the voice dialogue unit 20.
[0093] According to this configuration, the operator's voice is converted into text once, and therefore the operator's characteristics such as voice quality, volume, and emotion are removed from the second dialogue voice. By removing the operator's unique characteristics from the text on which the second dialogue voice is based, it becomes difficult for the user to distinguish whether the voice is based on the machine learning model 110 or the operator.
[0094] (4) The control unit 30 executes either the first dialogue process or the second dialogue process in accordance with the operator information relating to the state of the operator.
[0095] According to this configuration, either the first interaction process or the second interaction process can be executed depending on the operator information.
[0096] (5) The control unit 30 executes a first dialogue process when a voice dialogue is initiated in response to at least one of a user input action and the user's state, and executes a second dialogue process when a voice dialogue is initiated by operation of the operator terminal 200.
[0097] According to this configuration, when a voice dialogue is started by at least one of a user's input action and the user's state, the first dialogue process is executed, and therefore, a response to the user's request can be automatically made by the machine learning model 110. On the other hand, when a voice dialogue is started by an operation on the operator terminal 200, an operator can respond to the user's request.
[0098] (6) The control unit 30 executes either the first dialogue process or the second dialogue process in accordance with the user information relating to the attributes of the user.
[0099] According to this configuration, either the first interaction process or the second interaction process can be executed depending on the user information.
[0100] (7) The voice dialogue unit 20 has a display 25 that displays an avatar to the user. The display 25 displays an avatar that is common to the first dialogue process and the second dialogue process.
[0101] According to this configuration, a common avatar is displayed on the display 25 in the first dialogue process and the second dialogue process. Therefore, when the user has a voice dialogue with the voice dialogue system 10, it is difficult for the user to tell whether the voice dialogue is a dialogue with the machine learning model 110 or a dialogue with an operator.
[0102] <Modification> The above-described embodiment is an example of a form that the voice dialogue system 10 can take, and is not intended to limit the form. The voice dialogue system 10 can take a form different from the form exemplified in the above-described embodiment. Examples of such a form include a form in which part of the configuration of the above-described embodiment is replaced, changed, or omitted, or a form in which a new configuration is added to the embodiment. Modified examples of the embodiment are shown below.
[0103] <First Modification> In the voice dialogue system 10 according to the embodiment, when a voice dialogue with a user is started, whether to execute the first dialogue process or the second dialogue process is determined according to a predetermined selection condition. In the voice dialogue system 10 according to this modification, when a voice dialogue with a user is started, the first dialogue process is executed first. The voice dialogue system 10 according to the first modification will be specifically described below.
[0104] When starting a voice interaction, the control unit 30 first executes a first dialogue process. While the first dialogue process is being executed, the control unit 30 executes an interrupt process to start a second dialogue process in response to operations by the user and the operator. With this configuration, when a voice interaction with the user starts, the first dialogue process is executed by default first, and it is possible to quickly switch to the second dialogue process as needed.
[0105] The interrupt process is executed in response to at least one of a voice input by the user to the microphone 21 and an operation of the input device 23. The control unit 30 executes the interrupt process when a voice requesting the second dialogue process, such as "I would like to talk to an operator," is input from the user. The control unit 30 executes the interrupt process when the user inputs into the input device 23 that the user would like to execute the second dialogue process.
[0106] The interrupt process is executed in response to an operator's operation of the operator terminal 200. For example, the operator terminal 200 is configured so that the operator can check the history of the voice interaction in the first interaction process. During the execution of the first interaction process, if it is determined from the history of the voice interaction that the operator needs to directly respond to a user request, the interrupt process is executed in response to the operator's operation of the operator terminal 200.
[0107] The process in which the control unit 30 of this modified example executes either the first dialogue process or the second dialogue process will be described with reference to Fig. 6. After the process in Fig. 6 ends, the control unit 30 executes the process in Fig. 6 again after a predetermined time has elapsed.
[0108] In step S41, the control unit 30 determines whether or not to start a voice dialogue. If a voice dialogue is to be started, the control unit 30 proceeds to step S42. If a voice dialogue is not to be started, the control unit 30 ends the processing of FIG.
[0109] In step S42, the control unit 30 executes the first dialogue process, and then the process proceeds to step S43.
[0110] In step S43, the control unit 30 determines whether the voice dialogue has ended. If the voice dialogue has ended, the control unit 30 ends the first dialogue process and ends the process in Fig. 6. If the voice dialogue has not ended, the control unit 30 proceeds to step S44.
[0111] In step S44, the control unit 30 determines whether or not to execute interrupt processing. If the control unit 30 determines to execute interrupt processing, the process proceeds to step S45. If the control unit 30 determines not to execute interrupt processing, the process repeats from S42.
[0112] In step S45, the control unit 30 executes the second dialogue process, and then the process proceeds to step S46. The control unit 30 may end the first dialogue process when executing the second dialogue process.
[0113] In step S46, the control unit 30 determines whether the voice interaction has ended. If the voice interaction has ended, the control unit 30 ends the processing in Fig. 6. If the voice interaction has not ended, the control unit 30 repeats the processing of step S45.
[0114] <Second Modification> The control unit 30 may generate a re-learning dataset for the machine learning model 110 based on the second sentence information. The control unit 30 may update the machine learning model 110 by having the machine learning model 110 re-learn the re-learning dataset. According to this configuration, the machine learning model 110 is updated by the re-learning dataset generated based on the second sentence information, so that the text generated by the machine learning model 110 becomes closer to the operator's response. This reduces the sense of discomfort felt by the user in the dialogue between the voice dialogue system 10 and the user.
[0115] The flow of how the control unit 30 updates the machine learning model 110 will be described with reference to Fig. 7. In the flow of Fig. 7, steps S51 and S52 are added to the flow of Fig. 4. Explanation of parts common to the flow of Fig. 4 of the embodiment will be omitted.
[0116] After the process of step S25, the control unit 30 executes the process of step S51. In step S51, the control unit 30 generates a re-learning dataset from the second sentence information. The re-learning dataset is, for example, a set of information related to user input and information related to the operator's response. The information related to user input includes, for example, user text. The information related to user input may include user attributes. The information related to the operator's response includes, for example, response text based on the operator's response. The information related to the operator's response may include information related to the number of characters, number of words, tone of voice, ending, first person, etc. of the response text.
[0117] In step S52, the control unit 30 retrains the machine learning model 110 using the retraining dataset. Steps S51 and S52 may be performed each time the second dialogue process is executed, or may be performed when a predetermined number or more of information about user inputs and information about operator responses have been accumulated. Steps S51 and S52 may be performed before step S25, as long as they are performed after step S24.
[0118] <Third Modification> In the voice dialogue system 10 according to the embodiment, the control unit 30 forms a response text by converting the voice of the operator into text in the second dialogue processing. Then, the control unit 30 outputs second sentence information including the response text to the voice dialogue unit 20. In contrast, the control unit 30 may further convert the response text formed by converting the voice of the operator in the second dialogue processing based on sentence parameters. In the second dialogue processing, the control unit 30 outputs second sentence information including text obtained by converting the response text related to the operator's answer based on the sentence parameters.
[0119] According to this configuration, the text converted from the operator's voice is further converted based on the sentence parameters, so that the characteristics of the text included in the second sentence information are closer to the characteristics of the response text included in the first sentence information, thereby reducing the sense of discomfort felt by the user during the dialogue between the voice dialogue system 10 and the user.
[0120] As described above, the sentence parameters may be set with elements appropriate for when a responder speaks to a user at a facility, for example. If appropriate elements are set in the sentence parameters, even if the operator makes an inappropriate remark using a rude tone or the like, the response text regarding the inappropriate remark will be converted into appropriate text by the sentence parameters. As a result, the inclusion of inappropriate remarks in the text of the second sentence information is suppressed, and the quality of the voice (second dialogue voice) output based on the operator's remarks can be made uniform regardless of the personality or quality of the operator.
[0121] The flow in which the control unit 30 converts a response text related to an operator's answer based on sentence parameters will be described with reference to Fig. 8. In the flow in Fig. 8, step S61 is added to the flow in Fig. 4. Explanation of parts common to the flow in Fig. 4 of the embodiment will be omitted.
[0122] The control unit 30 inputs the response text regarding the operator's answer generated in step S23 to the machine learning model 110. In step S61, the machine learning model 110 converts the response text regarding the operator's answer based on the sentence parameters. The machine learning model 110 outputs the response text converted based on the sentence parameters to the control unit 30. The control unit 30 generates second sentence information including the response text converted based on the sentence parameters. The control unit 30 outputs the second sentence information to the voice dialogue unit 20.
[0123] <Fourth example of change> The control unit 30 may generate a database related to the results of the voice dialogue. The database stores data in which, for example, user text, sentence information including a response text to the user text, and dialogue voice generated from the sentence information are associated with each other.
[0124] For example, each time the control unit 30 outputs the first dialogue speech in the first dialogue processing, the control unit 30 forms first data in which user text, first sentence information, and the first dialogue speech output based on the first sentence information are associated with each other. The first data includes the user text generated from the user's speech in step S11, the first sentence information including the response text generated by the machine learning model 110 in step S13, and the first dialogue speech generated from the first sentence information in step S14.
[0125] For example, each time the control unit 30 outputs second dialogue speech in the second dialogue process, the control unit 30 forms second data in which user text, second sentence information, and second dialogue speech output based on the second sentence information are associated with each other. The second data includes user text generated from the user's speech in step S21, second sentence information including response text based on the operator's answer formed in step S23, and second dialogue speech generated from the second sentence information in step S24.
[0126] The control unit 30 saves the first data or the second data each time it forms the data. The control unit 30 generates a database by accumulating the data. The control unit 30 may output the first data or the second data to a computer such as another server each time it forms the data. With this configuration, the history of the voice dialogue performed by the machine learning model 110 and the history of the voice dialogue performed by the operator are generated as a database, so that the database can be used as a knowledge base.
[0127] <Other variations> The language used in the voice dialogue by the voice dialogue system 10 is not limited to Japanese. The voice dialogue system 10 may be able to change the language for the voice dialogue depending on the language used by the user.
[0128] In the voice dialogue system 10 according to the embodiment, the control unit 30 executes either a first dialogue process or a second dialogue process. Accordingly, either first text information or second text information is generated in response to a user request. Alternatively, the control unit 30 may generate both the first text information and the second text information in response to the user's voice. The control unit 30 complements either the first text information or the second text information based on the other of the first text information or the second text information. This configuration improves the accuracy of responses to the user because the text information includes both the text generated by the machine learning model 110 and the text related to the operator's response. The control unit 30 of this modification generates the second text information during execution of the first dialogue process. The control unit 30 of this modification generates the first text information during execution of the second dialogue process.
[0129] In the embodiment, the external server 100 has the machine learning model 110. Alternatively, the voice dialogue system 10 may have the machine learning model 110. The machine learning model 110 may be installed in a computer having the control unit 30.
[0130] In the embodiment, the voice dialogue unit 20 has the voice synthesis engine 26 and the sentence generation unit 27, but at least one of the voice synthesis engine 26 and the sentence generation unit 27 may be provided on the back-end side 12. A computer having the control unit 30 may have at least one of the voice synthesis engine 26 and the sentence generation unit 27. In this modification, the voice dialogue unit 20 and the control unit 30 communicate voice data such as the user's voice and dialogue voice.
[0131] In the voice dialogue system 10 according to the embodiment, the user's request is displayed as text on the display unit 220 of the operator terminal 200. Alternatively, the user's request may be output as voice from the communication device 210 of the operator terminal 200.
[0132] In the voice dialogue system 10 according to the embodiment, the operator inputs a response to a user request by voice at the operator terminal 200. Alternatively, the operator may input a response to a user request as text at the input device 23 at the operator terminal 200.
[0133] The display 25 may be omitted from the voice dialogue unit 20. In this modification, the user can also recognize the avatar through the voice generated by the voice synthesis engine 26. [Explanation of symbols]
[0134] 10...voice dialogue system, 20...voice dialogue unit, 21...microphone, 22...camera, 23...input device, 24...speaker, 25...display, 26...speech synthesis engine, 27...sentence generation unit, 30...control unit, 31...memory unit, 100...external server, 110...machine learning model, 200...operator terminal, 210...communication device, 220...display unit, 230...input unit.
Claims
1. A voice dialogue system, a voice dialogue unit that outputs to a user a dialogue voice generated by a voice synthesis engine that generates a dialogue voice based on specified voice parameters from text information including a response text; a control unit that executes a first dialogue process that causes the voice dialogue unit to output a first dialogue voice generated from first sentence information, and a second dialogue process that causes the voice dialogue unit to output a second dialogue voice generated from second sentence information, the first text information includes a response text regarding a response to the user that is generated by a machine learning model that generates text information by natural language processing; the second text information includes a response text regarding a response by the operator to the user, the response text being input to an operator terminal operated by the operator; The control unit executes either the first dialogue process or the second dialogue process in accordance with a predetermined selection condition.
2. 2. The voice dialogue system according to claim 1, wherein, in the first dialogue processing, the control unit causes the machine learning model to generate original text regarding a response to the user request, causes the machine learning model to form response text by converting the original text based on sentence parameters regarding sentence structure, forms the first sentence information including the response text based on the original text, and outputs the first sentence information to the voice dialogue unit.
3. The operator's voice is input to the operator terminal, 2. The voice dialogue system according to claim 1, wherein in the second dialogue processing, the control unit forms a response text by converting the operator's voice into text, forms the second sentence information including the response text based on the operator's voice, and outputs the second sentence information to the voice dialogue unit.
4. The voice dialogue system according to claim 1 , wherein the control unit executes either the first dialogue process or the second dialogue process in accordance with operator information relating to a state of the operator.
5. 4. The voice dialogue system according to claim 1, wherein the control unit executes the first dialogue processing when a voice dialogue is started in response to at least one of an input operation by the user and a state of the user, and executes the second dialogue processing when a voice dialogue is started by an operation of the operator terminal.
6. The voice dialogue system according to claim 1 , wherein the control unit executes either the first dialogue process or the second dialogue process in accordance with user information relating to attributes of the user.
7. the voice dialogue unit has a display that displays an avatar to the user; The voice dialogue system according to claim 1 , wherein the display displays the avatar that is common to the first dialogue process and the second dialogue process.
8. 4. The voice dialogue system according to claim 1, wherein the control unit generates a re-learning dataset for the machine learning model based on the second sentence information, and updates the machine learning model by re-learning the re-learning dataset to the machine learning model.
Citation Information
Patent Citations
Two-way video communication system and kiosk terminal
JP2019149630A