Information processing device, information processing program, and information processing method
The system addresses inefficiencies in automated response systems by automatically generating questions and using icebreaker utterances to manage user interactions, ensuring efficient operator response and reduced waiting times.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ADVANCE CREATE
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-26
AI Technical Summary
Existing systems face inefficiencies in operator response times due to variable automated response durations, leading to prolonged waiting times for users and reduced operator productivity.
Implementing a system that automatically responds to users for a predetermined time, generating questions, receiving answers, and determining the inclusion of icebreaker utterances based on remaining time and question count to efficiently transition to operator interaction.
This approach ensures operators can respond efficiently without extended user waiting times by fixing interview durations and utilizing icebreaker utterances to adjust for varying user response speeds.
Smart Images

Figure 2026086029000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing program, and an information processing method. In particular, for example, before an operator responds to a user, the present invention relates to an information processing apparatus, an information processing program, and an information processing method that automatically respond to the user.
Background Art
[0002] An example of an information processing system in the background art is disclosed in Patent Document 1. The automatic trading device operation support system of Patent Document 1 switches to the operator response mode after the automatic response mode by chat. In the operator response mode, the support terminal device receives history data and displays it on the support operation screen.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For example, if an operator responds to all inquiries from a user, the burden on the operator is large. Therefore, as in the above background art, it is conceivable to reduce the burden on the operator by automatically responding by chat. Also, in the above background art, since an automatic response can be made by chat until the operator becomes able to respond, the waiting time of the user can be reduced compared to responding only by the operator.
[0005] However, with the background technology described above, if the automated chat response mode cannot be used to respond, the system switches to operator response mode. Therefore, the length of time spent in automated response mode varies from user to user. As a result, if the automated response mode is used for a long time, operators may experience waiting times, which could prevent them from efficiently assisting users.
[0006] Therefore, it might be possible to fix the length of time spent in automated response mode for each user. However, if the automated response mode ends too quickly, the user will experience waiting time before the system switches to operator-assisted mode. Thus, there is room for improvement to prevent waiting time for both operators and users.
[0007] Therefore, the primary objective of this invention is to provide a novel information processing device, an information processing program, and an information processing method.
[0008] Another object of this invention is to provide an information processing device, an information processing program, and an information processing method that can efficiently respond to an operator without making the user wait as long as possible. [Means for solving the problem]
[0009] The first invention is an information processing device that automatically responds to a user for a predetermined time prior to an operator responding to the user, and includes a question generation means for sequentially generating a plurality of questions for the user, a question transmission means for transmitting the questions generated by the question generation means to a user terminal used by the user, an answer receiving means for receiving answers to the questions transmitted by the user terminal, a remaining time acquisition means for acquiring the remaining time after the predetermined time has been acquired by the answer receiving means, a determination means for determining whether or not to include an icebreaker based on the remaining time acquired by the remaining time acquisition means and the remaining number of items to be asked about, and the question generation means generates the next question including the icebreaker based on the icebreaker based if the determination means determines that an icebreaker based on the icebreaker should be included.
[0010] The second invention is subordinate to the first invention, and the question generation means generates questions concerning items other than those confirmed by the answers received by the answer receiving means, out of a plurality of items that the user should confirm.
[0011] The third invention is subordinate to the second invention, and the determination means determines whether to include an icebreaker-based utterance if, at a minimum, the remaining time is longer than the time required to ask questions relating to the remaining items and to obtain answers to those questions.
[0012] The fourth invention is an information processing program executed on an information processing device that automatically responds to a user for a predetermined time prior to an operator responding to the user, and the information processing device's processor performs the following steps: a question generation step of sequentially generating a plurality of questions for the user; a question transmission step of sending the questions generated in the question generation step to a user terminal used by the user; a response receiving step of receiving answers to the questions transmitted by the user terminal; a remaining time acquisition step of acquiring the remaining time after receiving answers in the response receiving step; a determination step of deciding whether or not to include an icebreaker based on the remaining time acquired in the remaining time acquisition step and the remaining number of items to ask about; and the question generation step generates the next question including the icebreaker based on the icebreaker if it is determined in the determination step to include an icebreaker based on the icebreaker.
[0013] The fifth invention is an information processing method for an information processing device that automatically responds to a user for a predetermined time prior to an operator responding to the user, wherein the processor of the information processing device sequentially generates a plurality of questions for the user, transmits the generated questions to a user terminal used by the user, receives answers to the questions transmitted by the user terminal, obtains the remaining time within the predetermined time when answers are received, determines whether to include an icebreaker based utterance based on the obtained remaining time and the number of remaining items to ask about, and if it is determined to include an icebreaker based utterance, generates the next question including the icebreaker based utterance. [Effects of the Invention]
[0014] According to this invention, during a predetermined time for an automated response system to respond automatically, the next question will be based on the remaining time and the number of remaining questions to ask the user. This allows the operator to respond efficiently without making the user wait for extended periods.
[0015] The above objects, other objects, features, and advantages of the present invention will become more apparent from the following detailed description of the embodiments made with reference to the drawings.
Brief Description of the Drawings
[0016] [Figure 1] FIG. 1 is a diagram showing an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of the electrical configuration of a central control device. [Figure 3] FIG. 3 is a block diagram showing an example of the electrical configuration of an operator terminal. [Figure 4] FIG. 4 is a block diagram showing an example of the electrical configuration of a user terminal. [Figure 5] FIG. 5 is a block diagram showing an example of the electrical configuration of an automatic response device. [Figure 6] FIG. 6 is a diagram showing an example of a response screen displayed on a display device of a user terminal. [Figure 7] FIG. 7 is a diagram showing an example of a hearing screen displayed on a display device of a user terminal. [Figure 8] FIG. 8 is a diagram showing an example of sequentially assigning users to respond to a certain operator. [Figure 9] FIG. 9 is a diagram showing a part of an example of a prompt input to a hearing execution model. [Figure 10] FIG. 10 is another part of an example of a prompt input to a hearing execution model, showing a part following the part of the prompt shown in FIG. 9. [Figure 11] FIG. 11 is a diagram showing an example of arranging questions, answers to these questions, and remaining time in a time series in a hearing. [Figure 12] FIG. 12 is a diagram showing another example of arranging questions, answers to these questions, and remaining time in a time series in a hearing. [Figure 13] ,FIG. 13 is a diagram showing an example of a memory map of the RAM of a central control device. [Figure 14]FIG. 14 is a diagram showing an example of a memory map of the RAM of the operator terminal. [Figure 15] FIG. 15 is a diagram showing an example of a memory map of the RAM of the user terminal. [Figure 16] FIG. 16 is a diagram showing an example of a memory map of the RAM of the automatic response device. [Figure 17] FIG. 17 is a flowchart showing a part of an example of the operator management process of the CPU of the central control device. [Figure 18] FIG. 18 is a flowchart that is another part of an example of the operator management process of the CPU of the central control device and follows FIG. 17. [Figure 19] FIG. 19 is a flowchart showing a part of an example of the response process of the CPU of the operator terminal. [Figure 20] FIG. 20 is a flowchart that is another part of an example of the response process of the CPU of the operator terminal and follows FIG. 19. [Figure 21] FIG. 21 is a flowchart that is another part of an example of the response process of the CPU of the operator terminal and follows FIG. 20. [Figure 22] FIG. 22 is a flowchart showing a part of an example of the hearing process of the CPU of the automatic response device. [Figure 23] FIG. 23 is a flowchart that is another part of an example of the hearing process of the CPU of the automatic response device and follows FIG. 22.
Embodiments for Carrying Out the Invention
[0017] Referring to FIG. 1, an information processing system 10 (hereinafter simply referred to as "system 10"), which is an embodiment of this invention, includes a central control device 12. The central control device 12 is communicably connected to an operator terminal 16, a user terminal 18, an automatic response device 20, etc. via a network 14. Further, the system 10 includes a hearing execution model 22, and the hearing execution model 22 is communicably connected to the automatic response device 20. However, the hearing execution model 22 may be included in the automatic response device 20 or may be communicably connected via the network 14.
[0018] This system 10 is applied to a store or facility (hereinafter referred to as "store, etc."). As will be described in detail later, it provides online remote support (customer service) to users using the user terminal 18.
[0019] The central control unit 12 is an information processing device such as a server, and it comprehensively controls the entire system 10.
[0020] The operator terminal 16 is an information processing device such as a personal computer, used by a human operator who is a staff member of a store or similar establishment. This operator terminal 16 is installed in the store's office, a contact center, and the operator's home.
[0021] The user terminal 18 is an information processing device such as a personal computer used for making inquiries, automatically entering and exiting the premises, and making payments, and is used by users of the store or other establishment. This user terminal 18 is installed in appropriate locations such as the entrance and exit of the store or establishment, the information desk, and the payment counter.
[0022] The stores, etc., include insurance and asset consultation offices, hotels, train stations, bus terminals, airports, hospitals, government offices, libraries, banks, theme parks, and shopping malls. The stores, etc., also include designated websites on the internet. In this case, the user terminal 18 is an information processing device such as a personal computer, smartphone, or tablet used by the user.
[0023] The automated response device 20 is an information processing device such as a personal computer or server, and it conducts interviews with the user. In other words, the automated response device 20 asks the user questions and obtains answers from the user. Furthermore, the automated response device 20 has a chatbot function that can automatically respond to the user, that is, an automated response function that utilizes artificial intelligence (i.e., AI).
[0024] The hearing execution model 22 is a large-scale language model for the automated response device 20 to automatically respond to, or conduct a hearing with, the user, and can use a large-scale language model such as ChatGPT. In this embodiment, the hearing execution model 22 generates questions to obtain answers from the user regarding pre-set items.
[0025] Furthermore, although Figure 1 shows one operator terminal 16, one user terminal 18, and one automated response device 20, in reality, one or more operator terminals 16, one or more user terminals 18, and one or more automated response devices 20 are connected to the network 14. Basically, the number of operator terminals 16 installed is less than the number of user terminals 18 installed. The operator terminals 16 and user terminals 18, and the user terminals 18 and automated response devices 20, can communicate bidirectionally via the network 14 using the P2P (Peer to Peer) method.
[0026] Network 14 consists of an IP network (or IP network) including the Internet, and an access network (or access network) for accessing this IP network. The access network can include public telephone networks, mobile phone networks, wired LANs, wireless LANs, CATV (Cable Television), etc. Furthermore, WebRTC (Web Real-Time Communication) is used to realize bidirectional communication using the P2P method. For this purpose, appropriate elements such as a signaling server, STUN server, and TURN server (not shown) are provided, but since these are publicly known, a detailed explanation is omitted.
[0027] Figure 2 is a block diagram showing an example of the electrical configuration of the central control unit 12. As shown in Figure 2, the central control unit 12 includes a CPU 30. The CPU 30 is connected to a RAM 32, a communication interface (hereinafter referred to as "communication I / F") 34, and an input / output interface (hereinafter referred to as "input / output I / F") 36 via an internal bus.
[0028] The CPU 30 is a processor that controls the central control unit 12 and, by extension, the entire system 10. However, instead of the CPU 30, a System-on-a-Chip (SoC) that includes multiple functions such as CPU functionality and GPU (Graphics Processing Unit) functionality may be provided.
[0029] RAM32 is the main memory of the central control unit 12 and is used as the work area or buffer area of the CPU 30. Although not shown in the diagram, the central control unit 12 is also provided with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in the auxiliary storage devices.
[0030] The communication interface 34 is a wired interface that, under the control of the CPU 30, transmits and receives control signals and data between the operator terminal 16, user terminal 18, and external computers such as the automated response device 20, via the network 14. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 34.
[0031] Input devices such as keyboards and computer mice, and display devices such as liquid crystal displays are connected to the input / output interface 36 as appropriate. The input / output interface 36 outputs operation data (or operation information) received from the input devices to the CPU 30. It also outputs image data generated by the CPU 30 to the display device, causing the screen or image corresponding to the image data to be displayed on the display device. Note that the configuration of the central control unit 12 shown in Figure 2 is just an example and is not limited to this configuration.
[0032] Figure 3 is a block diagram showing an example of the electrical configuration of the operator terminal 16. As shown in Figure 3, the operator terminal 16 includes a CPU 40. The CPU 40 is connected to the RAM 42, communication I / F 44, and input / output I / F 46 via an internal bus.
[0033] The CPU 40 is a processor that controls the overall operation of the operator terminal 16. However, instead of the CPU 40, an SoC (System on a Chip) that includes multiple functions such as CPU and GPU functions may be provided.
[0034] RAM 42 is the main memory of the operator terminal 16 and is used as the work area or buffer area of the CPU 40. Although not shown in the diagram, the operator terminal 16 is also equipped with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in the auxiliary storage devices.
[0035] The communication interface 44 is a wired interface for transmitting and receiving control signals and data between the central control unit 12, user terminals 18, and external computers such as the automated response device 20, via the network 14 under the control of the CPU 40. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 44.
[0036] The input / output interface 46 is connected to an input device 48, a display device 50, a microphone 52, a speaker 54, and a camera 56, among others. The input device 48 is a keyboard, computer mouse, and touch panel, etc. The display device 50 is, for example, a liquid crystal display. The microphone 52 is an operator voice detection means for detecting the voice emitted by the operator using the operator terminal 16, and is provided in a manner that can detect the operator's voice. The speaker 54 is a user voice output means for outputting the user's voice, and is provided in a manner that allows the operator to hear the sound output from the speaker 54. A headset in which the microphone 52 and speaker 54 are integrated may also be used. The camera 56 is an operator shooting means for shooting the operator using the operator terminal 16, and is provided in a manner that can shoot the operator.
[0037] The input / output interface 46 outputs operation data (or operation information) received from the input device 48 to the CPU 40. The input / output interface 46 also outputs image data generated by the CPU 40 to the display device 50, causing the display device 50 to display a screen corresponding to the image data. However, image data received from an external computer (for example, the user terminal 18) may also be output by the CPU 40.
[0038] Furthermore, the input / output interface 46 converts the operator's voice detected by the microphone 52 into digital audio data and outputs it to the CPU 40, and converts the audio data output by the CPU 40 into an analog audio signal and outputs it from the speaker 54. The audio data output by the CPU 40 includes the user's voice data received from the user terminal 18.
[0039] Furthermore, the input / output interface 46 outputs data of an image (operator image) including the operator captured (detected) by the camera 56 to the CPU 40. The image of the operator captured by the camera 56 may be a still image or a moving image. Note that the configuration of the operator terminal 16 shown in Figure 3 is just an example and is not limited to this configuration.
[0040] Figure 4 is a block diagram showing an example of the electrical configuration of the user terminal 18 shown in Figure 1. As shown in Figure 4, the user terminal 18 includes a CPU 60. The CPU 60 is connected to the RAM 62, communication I / F 64, and input / output I / F 66 via an internal bus.
[0041] The CPU 60 is a processor that controls the overall operation of the user terminal 18. However, instead of the CPU 60, an SoC (System on a Chip) that includes multiple functions such as CPU and GPU functions may be provided.
[0042] RAM62 is the main memory of the user terminal 18 and is used as the work area or buffer area of the CPU 60. Although not shown in the diagram, the user terminal 18 is also equipped with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in this auxiliary storage device.
[0043] Communication I / F64 is a wired interface that, under the control of the CPU 60, transmits and receives control signals and data between the central control unit 12, the operator terminal 16, and external computers such as the automated response device 20, via the network 14. However, a wireless interface for connecting to a wireless LAN can also be used as the communication I / F64.
[0044] The input / output interface 66 is connected to the aforementioned input device 68, display device 70, microphone 72, speaker 74, and camera 76, etc. This input / output interface 66 outputs operation data (or operation information) input from the input device 68 to the CPU 60.
[0045] Furthermore, the input / output interface 66 outputs image data generated by the CPU 60 to the display device 70, causing the display device 70 to display a screen corresponding to the image data. However, image data received from an external computer (for example, the operator terminal 16 and the automated response device 20) may also be output by the CPU 60.
[0046] Furthermore, the input / output interface 66 converts the user's voice detected by the microphone 72 into digital audio data and outputs it to the CPU 60, and converts the audio data output by the CPU 60 into an analog audio signal and outputs it from the speaker 74. The audio data output by the CPU 60 includes the operator's voice data received from the operator terminal 16. In addition, if a question from the automated response device 20 is output as audio, that audio data is also included.
[0047] Furthermore, the input / output interface 66 outputs data of an image (user image) including the user captured by the camera 76 to the CPU 60. The user image captured by the camera 76 may be a still image or a moving image. Note that the configuration of the user terminal 18 shown in Figure 4 is just an example and is not limited to this configuration.
[0048] Figure 5 is a block diagram showing an example of the electrical configuration of the automated response device 20. As shown in Figure 5, the automated response device 20 includes a CPU 80. The CPU 80 is connected to the RAM 82, communication I / F 84, and input / output I / F 86 via an internal bus.
[0049] The CPU 80 is the processor that controls the overall operation of the automated response system 20. However, instead of the CPU 80, a System of Control (SoC) that includes multiple functions such as CPU and GPU functions may be provided.
[0050] RAM 82 is the main memory of the automated response device 20 and is used as the work area or buffer area of the CPU 80. Although not shown in the diagram, the automated response device 20 is also provided with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in this auxiliary storage device.
[0051] The communication interface 84 is a wired interface for sending and receiving control signals and data between the central control unit 12 and external computers such as user terminals 18 via the network 14, under the control of the CPU 80. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 84.
[0052] Input devices 88 and display devices 90 are connected to the input / output interface 86. The input devices 88 include keyboards, computer mice, and touch panels. The display device 90 is, for example, a liquid crystal display.
[0053] The input / output interface 86 outputs operation data (or operation information) received from the input device 88 to the CPU 80, and also outputs image data generated by the CPU 80 to the display device 90, causing the display device 90 to display a screen or image corresponding to the image data. Note that the configuration of the automated response device 20 shown in Figure 5 is just one example and is not limited to this configuration.
[0054] In such a system 10, as described above, an operator using the operator terminal 16 can provide remote support (or customer service) to a user using the user terminal 18. In this case, a user who wishes to receive remote support can use any user terminal 18 to perform a call operation to summon the operator.
[0055] The central control unit 12 is notified that this call operation has been performed. Upon receiving this notification (i.e., a response request), the central control unit 12 determines an operator who can respond, assigns the user who made the response request to an available operator, and instructs the operator terminal 16 of the determined operator to respond to the user. In this embodiment, the central control unit 12 connects the operator terminal 16 of the determined operator with the user terminal 18 that made the call operation, that is, enables bidirectional communication using the P2P method. In other words, the connection between the operator terminal 16 and the user terminal 18 is established. Therefore, the determined operator begins the remote response.
[0056] In this remote interaction, operation data entered by the input device 68 of the user terminal 18, or the operation result corresponding to the operation data, is transmitted to the operator terminal 16. The image of the operation result corresponding to the operation data received by the operator terminal 16, or the image of the operation result received by the operator terminal 16, is then displayed on the display device 50 of the operator terminal 16. In addition, user voice data detected by the microphone 72 of the user terminal 18 (hereinafter referred to as "user voice data") is also transmitted to the operator terminal 16. The user voice corresponding to the user voice data received by the operator terminal 16 is output from the speaker 54 of the operator terminal 16. Furthermore, user image data captured by the camera 76 of the user terminal 18 (hereinafter referred to as "user image data") can also be transmitted to the operator terminal 16. In this case, the user image corresponding to the user image data received by the operator terminal 16 is displayed on the display device 50 of the operator terminal 16.
[0057] In addition, during remote support, the support screen 100 is displayed on the display device 70 of the user terminal 18. Figure 6 shows an example of the support screen 100. As shown in Figure 6, the support screen 100 displays an avatar image (hereinafter referred to as "avatar image") 102, which is a proxy character of the operator in charge of remote support. This avatar image 102 performs actions in accordance with the operator image captured by the camera 56 of the operator terminal 16. To this end, the operator terminal 16 performs a motion detection process on the operator image and generates avatar image data for displaying the avatar image 102 which performs actions in accordance with the operator image. This avatar image data is transmitted to the user terminal 18. Upon receiving the avatar image data, the user terminal 18 generates display screen data for displaying the support screen 100 based on the avatar image data, and the support screen 100 is displayed on the display device 70 based on this display screen data.
[0058] However, the response screen 100 may display an image of the operator themselves instead of the avatar image 102. Alternatively, the operator terminal 16 may generate action data corresponding to the operator's image and send it to the user terminal 18, and the user terminal 18 may generate the avatar image data.
[0059] In addition, operator voice data detected by the microphone 52 of the operator terminal 16 (hereinafter referred to as "operator voice data") is also transmitted to the user terminal 18. The operator voice corresponding to the operator voice data received by the user terminal 18 is output from the speaker 74 of the user terminal 18. The sound output from the speaker 74 may be voice processed by a voice changer.
[0060] Although not shown in the diagram, the user terminal 18 is provided with a call button or an icon that functions as a call button on the input device 68 or display device 70, etc., for receiving the aforementioned call operation. In addition, the display device 70 of the user terminal 18 in standby mode displays a message screen that includes information that remote support is available and that instructs the user to operate the call button if they wish to receive such remote support.
[0061] As described above, in this system 10, when a response request is received from a user terminal 18, the operator determined by the central control unit 12 basically responds using the operator terminal 16. However, questions regarding matters or information that the operator should obtain from the user in advance are handled by the automated response device 20 on behalf of the operator. In addition, for some of the matters or information that the operator should obtain from the user in advance (for example, basic information described later), the user can also input the information as text in a predetermined format that is displayed when a response request is sent to the central control unit 12.
[0062] Therefore, in this embodiment, when the operator interacts with a user, the operator first greets the user and asks them to answer questions from the automated response device 20 (hereinafter sometimes referred to as the "introduction to the interaction"), and the operator then requests the central control device 12 to conduct an interview with the user.
[0063] The central control unit 12 then connects the automated response device 20 with the user terminal 18 used by the user being responded to by the operator, that is, it enables bidirectional communication using the P2P method. In other words, the connection between the automated response device 20 and the user terminal 18 is established. Therefore, the automated response device 20 conducts an interview with the user.
[0064] Figure 7 shows an example of a hearing screen 150 displayed on the display device 70 of the user terminal 18. In the example shown in Figure 7, the hearing screen 150 contains questions for the user, as well as the user's answers to those questions. On the hearing screen 150, a display frame 152 for displaying the text of the questions for the user and a display frame 154 for displaying the text of the user's answers are arranged vertically in chronological order. The display frame 152 is positioned to the left of the hearing screen 150, and the display frame 154 is positioned to the right of the hearing screen 150.
[0065] The questions for the user are matters or content that the user needs to confirm regarding the services provided by System 10, and multiple items are prepared in advance. Each question is generated by the hearing execution model 22, and a display frame 152 showing the string of the generated question is displayed on the hearing screen 150.
[0066] However, if the user has already entered information for some items according to a predetermined format, questions will be asked to confirm the pre-entered items, and then questions will be asked about the other items, excluding those pre-entered items. In such cases, the automated response device 20 receives the information for the items entered according to the predetermined format from the central control unit 12 when establishing a connection with the user terminal 18, and stores it in the RAM 82 as the user's response.
[0067] When a user answers a question from the automated response device 20 using the user terminal 18, a display frame 154 containing the string of the answer is displayed on the hearing screen 150. The user's answer is also sent to the automated response device 20. The automated response device 20 receives the user's answer and inputs it into the hearing execution model 22. Based on the input string of the answer, the hearing execution model 22 generates questions (i.e., the next questions) to sequentially obtain answers for items for which answers have not yet been obtained, and outputs them to the automated response device 20. The automated response device 20 receives the next questions output from the hearing execution model 22 and sends them to the user terminal 18. Therefore, on the user terminal 18, a display frame 152 containing the string of the question is displayed on the hearing screen 150.
[0068] Users may enter their answers as text or by voice. If the answer is entered by voice, the user terminal 18 or the automated response device 20 will process the voice of the answer using speech recognition and transcribe it into text.
[0069] Once the user interview is complete, the answers to each question are transmitted from the automated response device 20 to the operator terminal 16 via the central control unit 12.
[0070] Furthermore, once the user interview is complete, the central control unit 12 instructs the operator terminal 16 to begin (resume) the user interaction. Therefore, the operator of the operator terminal 16 resumes the user interaction. However, before resuming the user interaction, the operator refers to and becomes aware of all the answers to each question received from the central control unit 12. As part of the interaction after the interview (hereinafter sometimes referred to as the "second half of the interaction"), the operator responds to questions from the user, asks questions to the user in return, explains the services provided by the system 10, and, in the case of sales, makes offers and closes. In this embodiment, the time for the second half of the interaction is set to the average time (for example, 10 minutes) when multiple operators actually interact with the user.
[0071] As described above, since the automated response system 20 is used to conduct interviews with users, the operator can attend to other users while the automated response system 20 is conducting an interview with one user. This allows for efficient user interaction.
[0072] However, the time it takes for each user to answer all questions during a hearing varies. Therefore, if an operator is assisting another user while one user is being interviewed, and then resumes assisting the first user after they have finished, the operator may have to wait for the first user to finish answering all questions. Consequently, to allow operators to assist users more efficiently, it might be advisable to fix the time allotted for each hearing.
[0073] However, if the interview time is fixed, users who have finished answering all questions before the fixed interview time has elapsed will have to wait for the operator to begin (or resume) the interaction.
[0074] Therefore, in this embodiment, regardless of whether the user has answered all the questions, the hearing time is fixed (or set) to a first predetermined time (for example, 20 minutes), so that the time at which the operator begins the second half of the interaction with the user is determined to be the time at which the user's hearing ends.
[0075] Furthermore, in this embodiment, the hearing time is fixed, and during the hearing, the remaining time for the hearing and the number of remaining questions are taken into consideration. At the very least, if the remaining time is longer than the time it will take to ask questions about the remaining items and obtain answers, that is, if there is ample time, icebreaker-based utterances are added to the questions, so that the hearing for users who answer quickly can be completed within the first predetermined time or approximately within the first predetermined time. Here, an icebreaker refers to one of the communication methods used to ease the atmosphere or to facilitate relationship building with the other party. As a result, the time that users have to wait for the operator to start (or resume) their response can be eliminated or almost eliminated.
[0076] However, the first predetermined time for the hearing is set to be longer than the time for the latter half of the response (hereinafter referred to as the "second predetermined time"). Therefore, while conducting a hearing for a previously assigned user, the operator can perform the introductory part of the response for another user assigned afterward, and then start a hearing for that other user. Furthermore, while conducting a hearing for another user, the operator can perform the latter half of the response for the previously assigned user. Therefore, the operator can respond to users more efficiently.
[0077] In this embodiment, when a user is assigned to an operator, the time for the introduction to the interaction is set to a third predetermined time (for example, 3 minutes). In reality, when an operator performs the introduction to the interaction with a user, there may be some deviation from the third predetermined time, but not a significant deviation, so it is considered that this will not affect the interaction with that user or other users.
[0078] Figure 8 illustrates an example of time changes when multiple users are sequentially assigned to a single operator. In Figure 8, each rectangle (or strip) represents the response time for one user. "OP" indicates the operator's response. Furthermore, "Hearing" refers to the hearing (i.e., response) by the automated response device 20, as described above. Also, as described above, the introductory part of the response is performed before the hearing, and the latter part of the response is performed after the hearing.
[0079] As shown in Figure 8, when a user requests assistance, the central control unit 12 determines which operator is available to handle the request. Whether an operator is available is determined based on whether a user is assigned to that operator, and, if an operator is assigned to handle the request, whether they are currently conducting a hearing with that user. However, users whose assistance by an operator has already been completed are not included in the list of users assigned to that operator.
[0080] In other words, if no user is assigned to respond to an operator, another user who made a response request will be assigned to that operator. Even if a user is assigned to respond to an operator, if that user is currently conducting a hearing and the latter half of that user's hearing overlaps with the hearing time of another user who made a response request, that other user will be assigned to respond.
[0081] However, in this embodiment, since the introductory part of the response is performed for each user, it is necessary to ensure that the time at which the introductory part of the response is started for a newly assigned user does not overlap with the latter part of the response for a previously assigned user.
[0082] Therefore, as shown in Figure 8, if a user is assigned to a particular operator (referred to here as "the operator") and is conducting an interview with this first user, the operator can then assign a second user. Furthermore, if the operator is conducting an interview with the second user, they can then assign a third user. In other words, users can be assigned to the operator sequentially. Note that the three users assigned to the operator are all different users.
[0083] However, as described above, the second user is assigned to the operator such that the introductory part of the second user's interaction does not overlap with the latter part of the first user's interaction, and the time spent listening to the second user overlaps with the entire latter part of the first user's interaction. Furthermore, the third user is assigned to the operator such that the introductory part of the third user's interaction does not overlap with the latter part of the second user's interaction, and the time spent listening to the third user overlaps with the entire latter part of the second user's interaction.
[0084] Furthermore, because users are assigned to the operator under the constraints described above, if another user is already assigned, the operator may not be able to immediately begin assisting that user. For this reason, when assigning a user to an operator, a start time for each user's interaction (hereinafter referred to as the "start time") is also set. The start time is the time when the introductory part of each user's interaction begins. Note that when assigning a user to an operator, the time for executing the introductory part of the interaction, the hearing, and the latter part of the interaction is fixed, so when the start time for the introductory part of the interaction is set, the start time for the hearing and the start time for the latter part of the interaction are also set.
[0085] Specifically, if no user is assigned to the operator, when a user requests assistance, the system will set the operator to assist the user at the current time or a few seconds to a few minutes in the future. When assigning a second or subsequent user, the start time is set considering the first predetermined time for the interview with the previously assigned user and the second predetermined time for the latter half of the interaction.
[0086] Therefore, while conducting the interview with the first user, the operator executes the introductory part of the response for the second user, initiating the interview with the second user. If the interview with the first user ends while the interview with the second user is in progress, the operator executes the latter half of the response for the first user. Similarly, while conducting the interview with the second user, and after completing the latter half of the response for the first user, the operator executes the introductory part of the response for the third user, initiating the interview with the third user. If the interview with the second user ends while the interview with the third user is in progress, the operator executes the latter half of the response for the second user. The operator handles subsequent users in the same manner.
[0087] I will omit the explanation, but similarly, users are assigned to other operators, and the assigned users respond in order.
[0088] By fixing the interview time in this way, it is possible to know in advance when each user's interview will end, allowing a single operator to sequentially assign multiple users to assist with minimal waiting time. Therefore, users can be assigned more efficiently than when interview times are not fixed.
[0089] Information regarding user assignments in a time-series format, as shown in Figure 8 (hereinafter referred to as "assignment information"), is managed for each operator by the central control unit 12. Based on this assignment information, the central control unit 12 determines which operator will be assigned to the user who made the response request.
[0090] When a user is assigned to an operator, the time for both the introductory and latter parts of the interaction is fixed. However, in reality, the actual interaction may be longer or shorter than the fixed time. As mentioned above, when an operator performs the introductory part of the interaction with a user, it may deviate slightly from the third predetermined time. However, since the introductory part only involves the operator greeting the user, it will not deviate significantly. Therefore, in this embodiment, once the introductory part of the interaction is completed, the hearing begins according to the operator's instructions. On the other hand, when an operator performs the latter part of the interaction with a user, if no restrictions are imposed, it may exceed the second predetermined time. In this case, the interaction with users assigned later will be delayed. For this reason, in this embodiment, the latter part of the interaction is forcibly terminated when the second predetermined time has elapsed.
[0091] Furthermore, as mentioned above, if the interview time is fixed to the first predetermined time, some users may finish answering all questions earlier than the predetermined time. In this case, users who have finished the interview early will have to wait before the operator can respond, which can be boring for them.
[0092] Therefore, in this embodiment, each time an answer is given to a question, the remaining time until the end of the hearing is considered, and if there is ample time remaining, an icebreaker-based utterance is added in addition to the question. By adding an icebreaker-based utterance, the question becomes longer, and the time until an answer is obtained can be extended. Therefore, even for users who answer quickly, the hearing can be adjusted to end within the first predetermined time or approximately the first predetermined time.
[0093] In this embodiment, the hearing execution model 22 functions as an AI that generates questions for the user. The hearing execution model 22 also determines whether or not to include an icebreaker-based utterance, and if it determines that an icebreaker-based utterance should be included, it generates the content of the icebreaker-based utterance and inserts (i.e., adds) the generated icebreaker-based utterance into the generated question.
[0094] Specifically, prompts like those shown in Figures 9 and 10 are input to the hearing execution model 22. Here, we will describe the case where system 10 provides a service that offers advice on asset formation.
[0095] As shown in Figure 9, the prompt for conducting the interview reads: "You are an assistant to an asset building advisor. You will conduct a one-by-one interview about the client's situation. The items to be covered in the interview are listed in a paragraph titled [Hearing Items]." 1. In the initial user message, we will send you known information, so please verify that information one by one first. 2. Conduct one-on-one interviews to address any remaining unclear points. Repeat and confirm the information each time. 3. Once all interviews are complete, summarize the information in a bulleted list and output it. 4. During the listening process (except for the final output in step 3), please output all text in a format that is easy to synthesize into speech. Do not use symbols such as parentheses () or quotation marks (「」). 5. Each time the user inputs, you will be given the remaining time. Compare the remaining time with the number of remaining interview questions using the following rules. If the remaining time is greater, include an icebreaker-based utterance to expand on the user's response. a. If remaining time (minutes) - remaining number of interview questions ÷ 2 > 3 The instructions (i.e., commands) state, "Confirm the known information and conduct interviews in the following order regarding any unclear points," along with a list of major and minor topics to be covered in the interviews, and examples of icebreaker-based utterances.
[0096] However, the instructions also include a method (in this case, a formula) for determining whether or not to include an icebreaker-based utterance. In this embodiment, if formula a is satisfied, that is, if the value obtained by subtracting the number of remaining hearing items (hereinafter referred to as "remaining number") divided by 2 from the remaining hearing time (minutes) is greater than 3, then it is decided to include an icebreaker-based utterance. If formula a is not satisfied, it is decided not to include an icebreaker-based utterance.
[0097] Furthermore, the "User Message" indicated in the prompt is information, i.e., items (sub-items), that the user has entered in advance according to a predetermined format. At the start of the hearing, the automated response device 20 acquires this information from the central control unit 12 and inputs it into the hearing execution model 22. The "User Input" indicated in the prompt is the content that is sent from the user terminal 18 to the automated response device 20 during the hearing and input from the automated response device 20 into the hearing execution model 22. In other words, it is the user's answers to the questions asked to the user during the hearing. The "Number of Remaining Hearing Items" indicated in the prompt is the number of remaining items to ask about. Furthermore, the "User Response" indicated in the prompt is the content that is generated by the hearing execution model 22 during the hearing and sent from the automated response device 20 to the user terminal 18. In other words, it is the questions asked to the user during the hearing.
[0098] The main categories covered in the interview are basic information, occupation and income, asset status, investment experience and risk tolerance, life event planning, tax and legal matters, expenses and living costs, financial goals, and other requests, and are listed in the order in which the interview takes place.
[0099] Following the information in Figure 9, as shown in Figure 10, the sub-items of the hearing are listed for each major item. Specifically, "[Hearing Items] Basic information: Name, age / date of birth, gender, address / contact information, family structure (spouse, children, parents, etc.) Occupation and Income: Occupation, employer / self-employment status, annual income (salary, bonus, other sources of income), spouse's occupation and income, years of service at current workplace, future career plans and retirement plans. Asset status: Current savings (regular savings, time deposits, investment trusts, etc.), currently held financial assets (stocks, bonds, real estate, etc.), debt status (mortgage, loans, credit card balances, etc.), liquidity of held assets (whether they can be withdrawn at any time). Investment experience and risk tolerance: Current investment experience (yes / no, years of experience, amount invested), investment objectives (wealth building, retirement plan, children's education fund, etc.), risk tolerance (how much risk you can accept), investment style (short-term, medium-term, long-term) Future plans and goals: Future plans (marriage, children's education, retirement, etc.), long-term goals (buying a house, studying abroad, second career, etc.), retirement life plans (where you want to live, how you want to spend your time), inheritance plans. Tax and Legal Matters: Tax situation (income tax, local tax, inheritance tax, etc.), legal documents (wills, trust agreements, etc.), whether tax advice is needed. Expenses and living expenses: Monthly living expenses (rent / mortgage, food, utilities, communication costs, etc.), annual fixed expenses (insurance premiums, taxes, education expenses, etc.), expenses for hobbies and entertainment, and planned special expenses (marriage, childbirth, education, travel, home renovations, etc.) Financial planning: Short-term goals (travel, car purchase, home appliances, etc.), medium-term goals (house purchase, education funds, etc.), long-term goals Other requests: This section includes information on satisfaction with the current plan, future concerns, support and advice expected from advisors, and any other special notes (individual requests or special circumstances).
[0100] The information in parentheses for each sub-item is a specific example related to that sub-item. Therefore, if the sub-item is "family structure," it is determined that a response has been obtained for that sub-item if responses regarding the presence or absence of a spouse, children, parents, etc., are received.
[0101] Furthermore, the prompt includes "[Examples of icebreaker-based utterances]" after each item in [Hearing Items]. That's amazing! Your annual income is well above the average. "I see, so you have experience with stocks. By the way, which securities company did you use?" is written.
[0102] Figures 11 and 12 show examples of questions generated when the automated response device 20 automatically responds to a user using the hearing execution model 22 with the prompts shown in Figures 9 and 10, respectively, along with the user's answers to the questions and the remaining time of the hearing when the answers were obtained, arranged in chronological order.
[0103] Figures 11 and 12 both show examples of questions and answers for a given user. However, the example in Figure 11 is for a user who answers questions at a normal or relatively slow pace, while the example in Figure 12 is for a user who answers questions at a relatively fast pace. To clearly illustrate when icebreaker-based utterances are included in the questions, examples are shown for the same user, both when the user answers at a normal or relatively slow pace and when they answer at a relatively fast pace.
[0104] In the prompt example described above, there are 37 sub-items of information to be obtained from the user, and the first predetermined time for the interview is 20 minutes. Therefore, if it takes approximately 32 seconds to ask the user a question about one sub-item and receive an answer, it is considered that answers for almost all sub-items can be obtained within the first predetermined time. For this reason, as described above, formula a is set to determine whether or not to include an icebreaker-based utterance. Consequently, if the number of sub-items and / or the length of the first predetermined time are changed, formula a will also be changed accordingly.
[0105] Therefore, if it takes less than 32 seconds to ask a user a question about a particular sub-item and receive an answer—that is, if the user responds quickly—an icebreaker-based utterance is generated and added to the text of the next question. In other words, the question is expanded or lengthened as needed.
[0106] Furthermore, if it takes longer than 32 seconds to ask a user a question about a particular sub-item and receive an answer, that is, if the user's response is slow, it may not be possible to obtain answers for all sub-items within the first predetermined time. However, this does not mean that the operator will have to wait to start (or resume) the interaction. For items for which an answer could not be obtained during the initial hearing, the operator can simply ask the user questions in the latter half of the interaction to obtain the answer, so this is not considered a particular problem.
[0107] As shown in Figure 11, the first question Q1 is asked to confirm the information obtained from the user message. Specifically, question Q1 is: "Let me confirm the information. First, regarding your basic information, your name is Taro Yamada, you are 52 years old, male, your contact number is xxx-xxxx-xxxx (phone number), and you have a spouse and children, correct?" The user's answer A1 is "Yes," and the remaining time for the interview is 19 minutes. At this point, the number of remaining sub-items is 32 (37 total) minus the number of sub-items for which answers have been obtained (5), and the remaining time to decide whether to include an icebreaker is 19 minutes. Therefore, applying this to the left side of equation a in the prompt, we get 19 - 32 ÷ 2 = 3, which does not satisfy equation a, so it is not decided to include an icebreaker.
[0108] Note that in formula a, the remaining time for the hearing is given in minutes, so seconds are rounded down. The same applies to the other answers.
[0109] Having obtained answer A1, the next question Q2 is generated according to the prompt. At this point, no icebreaker-based utterance is generated. Specifically, question Q2 is, "Thank you. Next, could you tell me about your occupation and income? First, could you tell me about your occupation?" The user's answer A2 to this is, "I am a company employee," and the remaining time (minutes) for the interview is 18 minutes. At this point, the number of remaining sub-items is 31 (37 total sub-items minus the number of sub-items for which answers have been obtained, 6), and the remaining time (minutes) to decide whether or not to include an icebreaker-based utterance is 18 minutes. Therefore, applying this to the left side of formula a in the prompt, 18 - 31 ÷ 2 = 2.5, which does not satisfy formula a, so it is not decided to include an icebreaker-based utterance.
[0110] After obtaining answer A2, the next question Q3 is generated according to the prompt. At this point, no icebreaker-based utterance is generated. Specifically, question Q3 is, "Thank you. Next, could you tell me about your workplace and whether you are self-employed?" The user's answer A3 to this is, "I am not self-employed," and the remaining time (minutes) for the interview is 17 minutes. At this point, the number of remaining sub-items is 30 (37 total sub-items minus the number of sub-items for which answers have been obtained, 7), and the remaining time (minutes) to decide whether to include an icebreaker-based utterance is 17 minutes. Therefore, applying this to the left side of formula a in the prompt, 17 - 30 ÷ 2 = 2, which does not satisfy formula a, so it is not decided to include an icebreaker-based utterance.
[0111] After obtaining answer A3, the next question Q4 is generated according to the prompt. At this point, no icebreaker-based utterance is generated. Specifically, question Q4 is, "I understand. Mr. Yamada Taro is an employee of Company Y and is not self-employed. Next, could you tell me your annual income?" The user's answer A4 to this is, "6 million yen." At this point, the remaining time (minutes) for the hearing is 15 minutes. At this point, the number of remaining sub-items after subtracting the number of sub-items for which answers have been obtained (8) from the total number of sub-items (37) is 29, and the remaining time (minutes) to decide whether or not to include an icebreaker-based utterance is 15 minutes. Therefore, applying this to the left side of formula a in the prompt, 15 - 29 ÷ 2 = 0.5, which does not satisfy formula a, so it is not decided to include an icebreaker-based utterance.
[0112] After obtaining answer A4, the next question Q5 is generated according to the prompt. At this point, no icebreaker-based utterances are generated. Specifically, question Q5 is, "Thank you. Your annual income is 6 million yen. Next, could you tell me how many years you have been working at your current workplace?"
[0113] Although not shown in the diagram, the automated response device 20 continues interviewing the user until answers are obtained for all sub-items or until the first predetermined time has elapsed. Furthermore, once answers are obtained for all questions, i.e., all sub-items, the interview execution model 22 generates data indicating that the interview has ended (interview completion data 504e, described later). In this case, the interview completion data 504e is sent to the user terminal 18, and instead of the interview screen 150, the user terminal 18 displays, for example, a screen listing all the information about the user of the user terminal 18.
[0114] On the other hand, as shown in Figure 12, when the user answers the question quickly, the remaining time for the interview when answer A1 is obtained is 19 minutes, and the remaining time for the interview when answers A2-A4 are obtained is 18 minutes in all cases. Therefore, when answers A1-A3 are obtained, equation a is not satisfied in any case, but when answer A4 is obtained, equation a is satisfied. Therefore, when answer A4 is obtained and question Q5 is generated, an utterance based on the icebreaker is also generated. For this reason, as shown in Figure 12, question Q5 includes an utterance based on the icebreaker, such as "Thank you. Your annual income is 6 million yen. That's wonderful. You are well above the average annual income. Next, could you tell me how many years you have been working at your current workplace?" Of this question Q5, the utterance based on the icebreaker is "That's wonderful. You are well above the average annual income."
[0115] In this way, when there is ample time remaining, the questions sent from the automated response device 20 to the user terminal 18 can be expanded (or made longer) by incorporating icebreaker-based utterances into the questions generated by the automated response device 20. In other words, the time the user has to read the questions is increased, and as a result, the time from asking a question to receiving an answer, i.e., the time for a single interaction, can be extended. Therefore, even for users who answer quickly, by appropriately incorporating icebreaker-based utterances into the questions, it is possible to adjust the system so that it takes the first predetermined time to answer all questions regarding all items. This avoids the inconvenience of the user having to simply wait for the operator to start (or resume) the interaction after finishing the hearing earlier than the first predetermined time.
[0116] Figure 13 shows an example of the memory map 200 of the RAM 32 built into the central control unit 12. As shown in Figure 13, the RAM 32 includes a program storage area 202 and a data storage area 204. The program storage area 202 stores the information processing program executed by the central control unit 12. This information processing program of the central control unit 12 includes a main processing program 202a, a communication program 202b, and an operator management program 202c, among others.
[0117] The main processing program 202a is a program for executing the main routine of information processing in the central control unit 12.
[0118] The communication program 202b is a program for communicating (sending and receiving data, etc.) with external devices such as the operator terminal 16, the user terminal 18, and the automated response device 20.
[0119] The operator management program 202c manages assignment information for each operator (or each operator terminal 16), assigns a user (or user terminal 18) to each operator (or operator terminal 16) based on this assignment information, and is a program that causes each operator to perform the assigned user's response.
[0120] The remaining time notification program 202d is a program that, when an inquiry is received from the automated response device 20 regarding a user who is currently undergoing a hearing, obtains or calculates the remaining time for the hearing for that user based on the count value of the first predetermined time counter 204c and notifies the automated response device 20 of this remaining time notification program. However, when the remaining time notification program 202d is executed, the communication program 202b is also executed.
[0121] Although not shown in the diagram, the program storage area 202 also stores other programs for executing the information processing program of the central control unit 12.
[0122] The data storage area 204 stores operator management data 204a, hearing completion data 204b, a first predetermined time counter 204c, and a second predetermined time counter 204d, among other things.
[0123] Operator management data 204a is data containing management information for each operator. In addition to assignment information for each operator, it includes data on the status of each user assigned to that operator, such as whether the initial part of the interaction is being performed, the interview is being conducted, or the latter part of the interaction is being performed (hereinafter referred to as "interaction status"). Management information for users for whom the operator has finished interacting is deleted.
[0124] The hearing completion data 204b is a bulleted list of all the answers obtained during the hearing process, i.e., all the information about the user. It is received from the automated response device 20 and stored in a way that allows each user to be identified.
[0125] The first predetermined time counter 204c is provided for each operator, for each user being handled, and is a counter or timer for counting the first predetermined time for each user's hearing.
[0126] The second predetermined time counter 204d is provided for each operator and is a counter or timer for counting the second predetermined time for the latter half of the user interaction being handled.
[0127] Although not shown in the diagram, the data storage area 204 stores other data necessary for the central control unit 12 to perform information processing, and also contains other timers (counters) and flags necessary for performing information processing.
[0128] Figure 14 shows an example of the memory map 300 of the RAM 42 built into the operator terminal 16. As shown in Figure 14, the RAM 42 includes a program storage area 302 and a data storage area 304. The program storage area 302 stores the information processing program to be executed on the operator terminal 16. This information processing program for the operator terminal 16 includes a main processing program 302a, an operation detection program 302b, a communication program 302c, a display program 302d, a voice detection program 302e, a voice output program 302f, a shooting program 302g, and a remote response program 302h.
[0129] The main processing program 302a is a program for executing the main routine of information processing on the operator terminal 16.
[0130] The operation detection program 302b is a program for detecting operation data 304a input from the input device 48 in accordance with the operator's actions.
[0131] The communication program 302c is a program for communicating with external devices such as the central control unit 12 and the user terminal 18.
[0132] The display program 302d is a program for generating and outputting display screen data necessary to display various screens on the display device 50.
[0133] The voice detection program 302e is a program for detecting voice using the microphone 52.
[0134] The audio output program 302f is a program for outputting audio from speaker 54.
[0135] The shooting program 302g is a program for capturing images (still images or moving images) using the camera 56.
[0136] The remote response program 302h is a program that allows an operator to remotely respond to a user using the operator terminal 16. This remote response program 302h includes various programs to enable remote response, including an avatar image data generation program for generating avatar image data.
[0137] Although not shown in the diagram, the program storage area 302 also stores other programs necessary for executing the information processing program on the operator terminal 16.
[0138] Various types of data are stored in the data storage area 304. These types of data include operation data 304a, image generation data 304b, operator voice data 304c, operator image data 304d, avatar image data 304e, user terminal data 304f, and hearing completion data 304g.
[0139] Operation data 304a is data representing the operator's operation status with respect to the input device 48, as detected by the operation detection program 302b.
[0140] Image generation data 304b consists of data such as polygon data and texture data used to generate display screen data by the display program 302d.
[0141] Operator voice data 304c is the operator's voice data detected by microphone 52.
[0142] Operator image data 304d is image data of the operator captured by camera 56.
[0143] Avatar image data 304e is image data of an avatar generated based on the captured image data of the operator.
[0144] User terminal data 304f is data transmitted and received between the user terminal 18 and the user terminal 18. User terminal data 304f includes the user's voice data and user image data received from the user terminal 18, and the operator's voice data and avatar image data sent to the user terminal 18.
[0145] The hearing completion data 304g is a bulleted list of all the answers obtained during the hearing process, i.e., all the information about the user. It is received from the central control unit 12 and stored in a way that allows each user to be identified.
[0146] Although not shown in the diagram, the data storage area 304 stores other data necessary for the operator terminal 16 to perform information processing, and also contains timers (counters) and flags necessary for performing information processing.
[0147] Figure 15 shows an example of a memory map 400 of RAM 62 built into the user terminal 18. As shown in Figure 15, RAM 62 includes a program storage area 402 and a data storage area 404. The program storage area 402 stores an information processing program that is executed on the user terminal 18. This information processing program for the user terminal 18 includes a main processing program 402a, an operation detection program 402b, a communication program 402c, a display program 402d, a voice detection program 402e, a voice output program 402f, a shooting program 402g, an operator response program 402h, and an automatic response program 402i, etc.
[0148] The main processing program 402a is a program for executing the main routine of information processing for the user terminal 18.
[0149] The operation detection program 402b is a program for detecting operation data 404a input from the input device 68 in accordance with user operations.
[0150] The communication program 402c is a program for communicating with external devices such as the central control unit 12, the operator terminal 16, and the automated response device 20.
[0151] The display program 402d is a program for generating display screen data necessary to display various screens on the display device 70.
[0152] The voice detection program 402e is a program for detecting voice using the microphone 72.
[0153] The audio output program 402f is a program for outputting audio from speaker 74.
[0154] The shooting program 402g is a program for capturing images (still images or moving images) using the camera 76.
[0155] The operator response program 402h is a program that works in cooperation with the operator terminal 16 to provide remote support, and includes various programs to enable remote support.
[0156] The automated response program 402i is a program for performing AI-based automated responses in cooperation with the automated response device 20, and includes various programs for realizing automated responses. In this embodiment, a hearing is performed between the automated response program and the automated response device 20.
[0157] Although not shown in the diagram, the program storage area 402 also stores other programs for executing the information processing program of the user terminal 18.
[0158] Various types of data are stored in the data storage area 404. These types of data include operation data 404a, image generation data 404b, user voice data 404c, user image data 404d, operator terminal data 404e, and automated response device data 404f.
[0159] Operation data 404a is data representing the user's operation status with respect to the input device 68, as detected by the operation detection program 402b.
[0160] Image generation data 404b consists of data such as polygon data and texture data used to generate display screen data by the display program 402d.
[0161] User voice data 404c is the user's voice data detected by microphone 72.
[0162] User image data 404d is image data of the user captured by camera 76.
[0163] Operator terminal data 404e is data transmitted and received between the operator terminal 16 and the operator terminal 16. Operator terminal data 404e includes the operator's voice data and avatar image data received from the operator terminal 16, and the user's voice data and user image data sent to the operator terminal 16.
[0164] The automated response system data 404f is data transmitted and received between the automated response system 20 and the automated response system 20. The automated response system data 404f includes data such as the question received from the automated response system 20, and data such as the answer to be sent to the automated response system 20. However, the questioner data and answer data are also used when generating display image data for the hearing screen 150 according to the display program 402d.
[0165] Although not shown in the diagram, the data storage area 404 stores other data necessary for the user terminal 18 to perform information processing, and also contains timers (counters) and flags necessary for performing information processing.
[0166] Figure 16 shows an example of the memory map 500 of the RAM 82 built into the automated response device 20. As shown in Figure 16, the RAM 82 includes a program storage area 502 and a data storage area 504. The program storage area 502 stores the information processing program executed by the automated response device 20. This information processing program for the automated response device 20 includes a main processing program 502a, a communication program 502b, and a hearing program 502c, among others.
[0167] The main processing program 502a is a program for executing the main routine of information processing for the automated response device 20.
[0168] The communication program 502b is a program for communicating with external devices such as the central control unit 12 and the user terminal 18.
[0169] The hearing program 502c is a program that inputs prompt data 504a to the hearing execution model 22, acquires question data 504b generated by the hearing execution model 22 and sends it to the user terminal 18, and receives answer data 504c for the question from the user terminal 18 and inputs it to the hearing execution model 22. In addition, when the hearing program 502c receives answer data 504c from the user terminal 18, it also acquires the remaining hearing time data 504d from the central control unit 12 and inputs the remaining time data 504d along with the answer data 504c to the hearing execution model 22. However, when the hearing program 502c is executed, the communication program 502b is also executed.
[0170] Although not shown in the diagram, the program storage area 502 also stores other programs for executing the information processing program of the automated response device 20.
[0171] The data storage area 504 stores prompt data 504a, question data 504b, answer data 504c, remaining time data 504d, and hearing completion data 504e, among others.
[0172] The prompt data 504a is data about the prompt to be entered into the hearing execution model 22, and is pre-generated or set by the administrator of system 10 or the provider of the service provided by system 10.
[0173] Question data 504b is data about questions asked to the user during the hearing process, and is generated and acquired for each user by the hearing execution model 22.
[0174] The response data 504c is the user's response data received from the user terminal 18 and is stored in a way that allows for user identification. The response data 504c, along with the remaining time data 504d, is input into the hearing execution model 22 for each user.
[0175] The remaining time data 504d is the data obtained from the central control unit 12 when the user's response data is received from the user terminal 18, representing the remaining time for the interview with that user.
[0176] The hearing completion data 504e is data about the final output in the hearing process, specifically, data that lists all the answers obtained during the hearing process, i.e., all the information about the user, and is generated and acquired by the hearing execution model 22. This hearing completion data 504e is transmitted to the central control unit 12, and from the central control unit 12, it is transmitted to the operator terminal 16 of the operator who is handling the user.
[0177] Although not shown in the diagram, the data storage area 504 stores other data necessary for the automated response device 20 to perform information processing, and also contains timers (counters) and flags necessary for performing information processing.
[0178] Figures 17 and 18 are flowcharts showing an example of operator management processing, which is an example of information processing by the CPU 30 of the central control unit 12. Operator management processing is performed for each operator terminal 16, its operator, and the user assigned to that operator. The following describes the operator management processing when a certain operator terminal 16, its operator, and a certain user are assigned to that operator, but the same applies to other users assigned to that operator terminal 16 or its operator. The same also applies to other operator terminals 16, their operators, the users assigned to those operators, and other users. Furthermore, in this operator management processing, a certain operator terminal 16 will be referred to as "the operator terminal 16," its operator as "the operator," and a certain user assigned to that operator as "the user."
[0179] Although not shown in the diagram, when the CPU 30 of the central control unit 12 receives a response request from a user terminal 18, it determines which operator is available to respond to the user of that user terminal 18. The method for determining which operator is available is as described above. Once the CPU 30 has determined which operator is available, that is, has assigned the user of the user terminal 18 that made the response request to that operator, it executes operator management processing for that user in accordance with the start time of the response.
[0180] As shown in Figure 17, when the CPU 30 starts the operator management process, in step S1 it instructs the operator terminal 16 to respond to the user. In step S1, the CPU 30 updates the operator management data 204a. Specifically, the CPU 30 adds management information to the operator management data 204a, linked to the operator, which includes assignment information that assigns the user to the operator on the operator terminal 16, and indicates that the introductory part of the response is being executed for the user.
[0181] In the next step, S3, it is determined whether there is a request from the operator terminal 16 to start a hearing about the user. If the answer in step S3 is "NO," that is, if there is no request from the operator terminal 16 to start a hearing about the user, the process returns to step S3.
[0182] On the other hand, if the answer in step S3 is "YES," that is, if the operator terminal 16 requests the start of an interview about the user, in step S5, the automated response device 20 is instructed to conduct an interview about the user, in step S7, the countdown for the first predetermined time begins, and the process proceeds to step S9. Also in step S7, the CPU 30 updates the operator management data 204a by updating it with management information that includes assignment information assigning the user to the operator of the operator terminal 16 and indicates that an interview is currently underway with the user.
[0183] In step S9, it is determined whether the automated response system 20 has inquired about the remaining time for the interview with the user in question. If the answer in step S9 is "NO," that is, if the automated response system 20 has not inquired about the remaining time for the interview with the user in question, the process proceeds to step S13.
[0184] On the other hand, if the answer in step S9 is "YES," that is, if the automated response device 20 inquires about the remaining time for the hearing regarding the user, then in step S11, the remaining time is sent to the automated response device 20, and the process proceeds to step S13. In step S11, the CPU 30 calculates the remaining time by subtracting the count value of a counter that counts the first predetermined time from the first predetermined time. However, in this embodiment, seconds are rounded down.
[0185] In step S13, it is determined whether a first predetermined time (for example, 15 minutes) has elapsed since the start of the interview. If the answer in step S13 is "YES," that is, if the first predetermined time has elapsed since the start of the interview, in step S15, the automated response device 20 is instructed to end the interview for that user, and the process proceeds to step S17.
[0186] On the other hand, if the answer in step S13 is "NO," that is, if the first predetermined time has not elapsed since the start of the hearing, then in step S17, it is determined whether or not a notification of the end of the hearing has been received from the automated response device 20. In other words, the CPU 30 determines whether or not it has received the hearing completion data.
[0187] If the answer in step S17 is "NO," that is, if there is no notification from the automated response device 20 that the hearing has ended, the process returns to step S9. On the other hand, if the answer in step S17 is "YES," that is, if there is a notification from the automated response device 20 that the hearing has ended, the process proceeds to step S19 as shown in Figure 18. Although not shown in the diagram, if the answer in step S17 is "YES," the CPU 30 updates the operator management data 204a by updating it with management information that includes assignment information assigning the user to the operator on the operator terminal 16, and indicates that the latter half of the response is being executed for that user.
[0188] As shown in Figure 18, in step S19, the operator terminal 16 is called, and in step S21, the user's hearing completion data 204b is sent to the operator terminal 16, and the process proceeds to step S23.
[0189] In step S23, the operator terminal 16 is instructed to respond to the user; in step S25, the countdown for the second predetermined time begins; and in step S27, it is determined whether the user's response has ended. Here, the CPU 30 determines whether it has received notification from the operator terminal 16 that the response has ended, or whether the second predetermined time has elapsed.
[0190] If the answer in step S27 is "NO," that is, if the interaction with the user has not ended, the process returns to step S27. On the other hand, if the answer in step S27 is "YES," that is, if the interaction with the user has ended, the operator management process for that user is terminated. However, if the second predetermined time has elapsed, the CPU 30 instructs the operator terminal 16 to terminate the interaction. Furthermore, when the CPU 30 terminates the operator management process for that user, it updates the operator management data 204a by updating it with management information that includes the assignment information that assigned the user to the operator on the operator terminal 16 and indicates that the interaction with that user has ended. However, the management information for the user whose interaction has ended may be deleted.
[0191] Figures 19-21 are flowcharts showing an example of response processing included in the information processing of the CPU 40 of the operator terminal 16 shown in Figure 3. When the CPU 40 of the operator terminal 16 receives a response instruction from the central control unit 12, it starts response processing. However, the response processing is performed for each user to whom it is assigned to respond. Therefore, if multiple users are assigned to the operator, the response processing is performed in parallel for each user. In addition, although not shown in the diagram, processing to receive control signals or data from other devices and image capture processing are also performed in parallel with the response processing.
[0192] As shown in Figure 19, when the CPU 40 starts the response process, in step S101 it establishes a connection with the user terminal 18 (hereinafter referred to as "the user terminal 18" in this response process) used by the user who issued the response instruction (hereinafter referred to as "the user" in this response process). However, the connection information of the user terminal 18 is transmitted from the central control unit 12 to the operator terminal 16 along with the response instruction.
[0193] In the next step, S103, it is determined whether the operator's voice has been detected. If the answer in step S103 is "NO," that is, if the operator's voice has not been detected, the process proceeds to step S107.
[0194] On the other hand, if the answer in step S103 is "YES," that is, if the operator's voice has been detected, then in step S105, the operator's voice data is sent to the user terminal 18, and the process proceeds to step S107.
[0195] In step S107, avatar image data 304e is generated based on operator image data 304d, and in step S109, the avatar image data 304e is sent to the user terminal 18, and the process proceeds to step S111.
[0196] In step S111, it is determined whether data has been received from the user terminal 18. Here, the CPU 40 determines whether user voice data and / or user image data of the user has been received. If the answer in step S111 is "NO", that is, if no data has been received from the user terminal 18, the process proceeds to step S115.
[0197] On the other hand, if the answer in step S111 is "YES," that is, if data is received from the user terminal 18, then in step S13, the user's voice and / or the user's image are output, and the process proceeds to step S115.
[0198] Step S115 determines whether the operator has given the instruction to start the interview. If the answer in Step S115 is "NO," that is, if the operator has not given the instruction to start the interview, the process returns to Step S103.
[0199] On the other hand, if the answer in step S115 is "YES," that is, if the operator gives an instruction to start the hearing, then in step S117, a request to start the hearing is sent to the central control unit 12, and the process proceeds to step S119 shown in Figure 20. Steps S103 to S115 correspond to the introduction of the dialogue.
[0200] As shown in Figure 20, step S119 determines whether there is a call from the central control unit 12. If the answer in step S119 is "NO," that is, if there is no call from the central control unit 12, the process returns to step S119. On the other hand, if the answer in step S119 is "YES," that is, if there is a call from the central control unit 12, the process proceeds to step S121.
[0201] In step S121, it is determined whether or not the hearing completion data 304g has been received from the central control unit 12. If the answer in step S121 is "NO," that is, if the hearing completion data 304g has not been received from the central control unit 12, the process returns to step S121. On the other hand, if the answer in step S121 is "YES," that is, if the hearing completion data 304g has been received from the central control unit 12, the process proceeds to step S123.
[0202] In step S123, all information about the user is displayed in a bulleted list. Therefore, the operator can see the user's answers during the interview process and whether there are any unanswered questions. Accordingly, the operator refers to the user's answers during the interview process and performs the latter part of the interaction with the user.
[0203] As shown in Figure 21, the next step S125 determines whether the operator's voice has been detected. If the answer in step S125 is "NO", the process proceeds to step S129. On the other hand, if the answer in step S125 is "YES", the process proceeds to step S127, where the operator's voice data is sent to the user terminal 18, and the process proceeds to step S129.
[0204] In step S129, avatar image data 304e is generated based on operator image data 304d, and in step S131, the avatar image data 304e is sent to the user terminal 18, and the process proceeds to step S133.
[0205] In step S133, it is determined whether or not data has been received from the user terminal 18. If the answer in step S133 is "NO", the process proceeds to step S137. On the other hand, if the answer in step S133 is "YES", then in step S135, the user's voice and / or image are output, and the process proceeds to step S137.
[0206] In step S137, it is determined whether there is an instruction to end the interaction. Here, the CPU 40 determines whether the operator has entered an instruction to end the interaction, or whether it has received an instruction to end the interaction from the central control unit 12.
[0207] If the answer in step S137 is "NO," that is, if there is no instruction to end the interaction, the process returns to step S125. On the other hand, if the answer in step S137 is "YES," that is, if the operator gives an instruction to end the interaction, the interaction processing for that user is terminated. Steps S125-S137 correspond to the latter half of the dialogue. However, if the operator gives an instruction to end the interaction, the CPU 40 notifies the central control unit 12 of the termination of the interaction.
[0208] Figures 22 and 23 are flowcharts showing an example of the hearing process included in the information processing of the CPU 80 of the automated response device 20 shown in Figure 5. When the CPU 80 of the automated response device 20 receives a hearing instruction from the central control unit 12, it starts the hearing process. However, the hearing process is executed for each user that one or more operators are responding to. Therefore, when hearings are conducted for multiple users, the hearing process is executed in parallel for each user. The CPU 80 of the automated response device 20 has previously input the prompt data 504a into the hearing execution model 22.
[0209] As shown in Figure 22, when the CPU 80 starts the hearing process, in step S201 it establishes a connection with the user terminal 18 of the user being interviewed. However, the automated response device 20 receives from the central control unit 12, along with the instruction to conduct the hearing, the connection information of the user terminal 18 of the user being interviewed and the user message data that the user has previously entered into a format.
[0210] In the next step, S203, the remaining time for the hearing is obtained from the central control unit 12. Here, the CPU 80 queries the central control unit 12 for the remaining time for the user's hearing and obtains the remaining time transmitted from the central control unit 12.
[0211] In the next step, S205, the remaining time and response data are entered into the hearing execution model 22. At the start of the hearing process, the response data 504c is the user message data, that is, data about the information that was previously entered into the format.
[0212] Next, in step S207, it is determined whether or not data has been received from the hearing execution model 22. If the answer in step S207 is "NO," that is, if no data has been received from the hearing execution model 22, the process returns to step S207.
[0213] On the other hand, if the answer in step S207 is "YES," that is, if data has been received from the hearing execution model 22, then in step S209, it is determined whether the hearing has ended. Here, the CPU 80 determines whether the received data is hearing completion data 504e.
[0214] If the answer in step S209 is "NO," that is, if the hearing is not finished, then in step S211, the question data 504b is sent to the user terminal 18, and the process proceeds to step S215 as shown in Figure 23. Therefore, on the user terminal 18, a display frame 152 showing the question is displayed on the hearing screen 150.
[0215] On the other hand, if the answer in step S209 is "YES," that is, if the hearing is complete, then in step S213, hearing completion data 504e is sent to the user terminal 18, and the process proceeds to step S219 shown in Figure 23. Therefore, instead of the hearing screen 150, the user terminal 18 displays a screen listing all the information about the user.
[0216] As shown in Figure 23, in step S215, it is determined whether or not the response data 504c has been received from the user terminal 18. If the answer in step S215 is "YES", that is, if the response data 504c has been received from the user terminal 18, the process returns to step S203.
[0217] On the other hand, if the answer in step S215 is "NO," that is, if no response data 504c has been received from the user terminal 18, then in step S217, it is determined whether or not there is an instruction from the central control unit 12 to end the hearing. If the answer in step S217 is "NO," that is, if there is no instruction from the central control unit 12 to end the hearing, the process returns to step S215. On the other hand, if the answer in step S217 is "YES," that is, if there is an instruction from the central control unit 12 to end the hearing, then in step S219, hearing completion data is sent to the central control unit 12 to end the hearing process. However, the hearing completion data sent in step S219 is data for one or more answers to one or more questions in the hearing, and does not necessarily contain all answers to all questions.
[0218] According to this embodiment, if there is sufficient time remaining for the automated response system to hear the user, icebreaker-based utterances are included in the questions as appropriate, so that the hearing can be completed within the first predetermined time or approximately the first predetermined time, and there is no inconvenience of making the user wait until the operator begins to respond. In other words, the operator can respond to the user with as little waiting time as possible.
[0219] Furthermore, according to this embodiment, since the time for the automated response device to hear from the user is fixed to a first predetermined time, the time for the operator to respond to that user afterward can be predicted. Therefore, if a response request comes in from another user, the operator can be efficiently assigned to respond to that other user. Thus, the operator can efficiently respond to each of multiple users.
[0220] In the above embodiment, the central control unit 12 and the automated response device 20 are provided separately, but the central control unit 12 and the automated response device 20 can also be used interchangeably.
[0221] Furthermore, although the above embodiment includes an introduction section for handling customer interactions, it is not necessary to include such an introduction section.
[0222] Furthermore, in the above-described embodiment, the operator responded using an avatar and voice, but the response could also be made via text chat. In addition, if only voice responses are used, the system could be configured such that the central control unit, operator terminal, user terminal, and automated response device each make calls or communicate via the telephone network.
[0223] Furthermore, while the above-described embodiment explained the case where a CG avatar, acting as a proxy character for the operator, is used for communication, a robot avatar can also be used. In such a case, a robot is provided in place of the user terminal, and the robot's movements and speech are controlled by the operator or automatically. The robot also acquires the user's image and voice and transmits them to the operator terminal or automated response device. In other words, the term "avatar" encompasses remotely controlled robots and CG agents, and should be understood to include robots and CG agents that interact with the user partially or completely autonomously. When the user terminal is a robot, it is equipped with movable parts such as arms, and the operator's movements are represented by the operation of these movable parts.
[0224] Furthermore, in the above-described embodiment, the automated response device conducts the interview using an AI-powered chatbot. However, the interview may also be conducted by sending all questions to the user's terminal in a predetermined format or by making a call using voice guidance. In such cases as well, depending on the remaining time for the interview, spoken text or audio based on icebreakers may be appropriately inserted into the questions.
[0225] Furthermore, in the above embodiment, when assigning a user to an operator, the latter half of the interaction is set to a second predetermined time, and when responding to a user, the interaction does not exceed the second predetermined time. However, when responding to a user, it may be permitted to exceed the second predetermined time. In such a case, the start time of the latter half of the interaction for the subsequently assigned user is delayed. If the start time of the latter half of the interaction is delayed, a video explaining the services provided by the system may be played from the end of the hearing until the start time of the latter half of the interaction.
[0226] Furthermore, the order in which each step in the flowchart shown in the above embodiment is processed can be changed if the same result can be obtained. Also, the various screens and specific numerical values shown in this specification and drawings are merely examples and can be changed as needed. [Explanation of Symbols]
[0227] 10. Information Processing Systems 12 ... Central Control Unit 14…Network 16 ... Operator terminal 18 ... User terminal 20 ...Automated response system 22...Hearing Implementation Model
Claims
1. An information processing device that automatically responds to a user for a predetermined period of time prior to the operator responding to the user, Question generation means for sequentially generating multiple questions for the user, Question transmission means that transmits the questions generated by the question generation means to the user terminal used by the user. Answer receiving means for receiving the answer to the question transmitted by the user terminal, When a response is received by the response receiving means, the remaining time acquisition means acquires the remaining time of the predetermined time. A determination means for determining whether or not to include an icebreaker based on the remaining time obtained by the remaining time acquisition means and the remaining number of questions to be asked. The question generation means is an information processing device that generates the next question including an icebreaker-based utterance when the determination means determines that an icebreaker-based utterance should be included.
2. The information processing apparatus according to claim 1, wherein the question generation means generates questions concerning items other than those confirmed by the answers received by the answer receiving means, out of a plurality of items to be confirmed by the user.
3. The information processing apparatus according to claim 2, wherein the determination means determines to include an icebreaker-based utterance if the remaining time is longer than the time required to ask the questions relating to the remaining number of items and to obtain the answers to those questions.
4. An information processing program executed on an information processing device that automatically responds to the user for a predetermined time prior to the operator responding to the user, The processor of the aforementioned information processing device, A question generation step that sequentially generates multiple questions for the user, A question transmission step in which the question generated in the question generation step is sent to the user terminal used by the user, A response receiving step, which receives the answer to the question transmitted by the user terminal, When a response is received in the response receiving step, a remaining time acquisition step is taken to acquire the remaining time of the predetermined time. A decision step to determine whether or not to include an icebreaker based on the remaining time obtained in the remaining time acquisition step, and the remaining number of questions to be asked. The question generation step is an information processing program that, if it is determined in the decision step to include an icebreaker-based utterance, generates the next question including the icebreaker-based utterance.
5. An information processing method for an information processing device, which automatically responds to the user for a predetermined period of time prior to the operator responding to the user, The processor of the aforementioned information processing device is Multiple questions are sequentially generated for the aforementioned user, The generated question is sent to the user terminal used by the user. The user receives the answer to the question transmitted by the user terminal, Upon receiving the aforementioned response, obtain the remaining time of the predetermined period, Based on the remaining time obtained and the number of remaining questions, decide whether or not to include an icebreaker-based utterance. An information processing method that, when it is determined to include an utterance based on the icebreaker, generates the next question that includes the utterance based on the icebreaker.