Survey processing system, survey processing method, and survey processing program
The system addresses high costs and accuracy issues in conventional survey systems by using speech recognition and evaluation to convert and aggregate answers, enabling efficient and accurate survey results on operator interactions.
Patent Information
- Application Number
- JP2024029726
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2026-03-02
- Estimated Expiration
- 2044-02-29
AI Technical Summary
Conventional survey systems incur high costs due to the need for personalized AI learning for each customer and are unable to accurately capture responses to specific events, such as operator interactions during calls, placing a burden on both operators and customers.
A system that utilizes speech recognition to convert conversational voice into text and outputs answers to questionnaire items, incorporating a speech recognition evaluation to ensure high accuracy and a configuration that allows for prompt conversion and answer aggregation, with the ability to issue alerts when necessary.
Enables the acquisition of survey results regarding operator responses with a simple configuration, ensuring accurate and efficient processing of customer interactions, reducing psychological anxiety for operators, and facilitating easy understanding of customer impressions.
Smart Images

Figure 0007821972000001 
Figure 0007821972000002 
Figure 0007821972000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a questionnaire processing system, a questionnaire processing method, and a questionnaire processing program for obtaining the results of a questionnaire. [Background technology]
[0002] In the past, managers of call centers and the like would conduct a survey of customers who called to determine whether an operator's response to a customer inquiry was appropriate. Generally, when conducting a survey, responses to the survey are obtained from multiple customers and the responses are tallied.
[0003] Furthermore, a system is known that uses multiple personalized AIs (Artificial Intelligences) to conduct surveys in order to reduce the workload of survey research (see Patent Document 1). With this conventional system, it is possible to obtain answers from personalized AIs to questions such as, "Will you support Mr. X in the next presidential election?" [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7101357 Summary of the Invention [Problem to be solved by the invention]
[0005] For example, if one were to obtain responses to a customer questionnaire using the conventional system described in Patent Document 1, it would be necessary to have the AI learn the personality of each person (i.e., each customer), which would incur a very large cost.
[0006] Furthermore, with conventional systems, it is not possible to expect accurate responses to surveys about specific events that are not included in the learning data, such as the responses of operators (i.e., the person who answered) to actual customer inquiries over the phone (for example, a specific call recorded on a specific date at a specific time). On the other hand, it is unrealistic to directly survey customers who called immediately after they ended the call, and this places a heavy burden on both the operator and the customer.
[0007] Therefore, the main purpose of the present disclosure is to provide a survey processing system, a survey processing method, and a survey processing program that can obtain, with a simple configuration, the results of a survey regarding services provided through two-way voice calls between the service provider and the customer, including the response of an operator who speaks with the customer. [Means for solving the problem]
[0008] The questionnaire processing system of the present disclosure includes: Regarding the subject conversation between the customer and the operator, from the perspective of the customer who is a participant in said conversation Survey of answer A questionnaire processing system for acquiring the questionnaire answer the one or more processors execute a process for acquiring questionnaire items as questions related to the questionnaire; Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversation By instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The answer output unit, which is capable of outputting answers to questions posed to voice, is instructed to output answers to the questionnaire items in the voice recognition results, and the answers to the questionnaire items are obtained from the answer output unit.
[0009] In addition, the questionnaire processing method of the present disclosure includes: Regarding the subject conversation between the customer and the operator, from the perspective of the customer who is a participant in said conversation Survey of answera questionnaire processing method for acquiring a questionnaire item, the method comprising: one or more computers acquiring questionnaire items as questions related to the questionnaire; Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversation By instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The answer output unit, which is capable of outputting answers to questions posed to voice, is instructed to output answers to the questionnaire items in the voice recognition results, and the answers to the questionnaire items are obtained from the answer output unit.
[0010] In addition, the questionnaire processing program of the present disclosure Between the customer and the operator Target conversations , from the perspective of a customer who is a participant in the conversation Survey of answer It is a questionnaire processing program that causes a computer to perform information processing to obtain And, in front The information processing includes acquiring questionnaire items as questions regarding the questionnaire, Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversation By instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The configuration includes a procedure for acquiring answers to the questionnaire items from an answer output unit capable of outputting answers to questions asked to voice by instructing the answer output unit to output answers to the questionnaire items in the voice recognition results. [Effects of the Invention]
[0011] According to the present disclosure, the results of a survey regarding services provided by a service provider through two-way voice calls with a customer, such as the response of an operator who spoke with the customer, can be obtained with a simple configuration. [Brief explanation of the drawings]
[0012] [Figure 1] Overall configuration diagram of a questionnaire processing system according to the first embodiment [Figure 2] A block diagram showing the schematic configuration of the operator terminal shown in Figure 1. [Figure 3] FIG. 1 is a flow chart showing the flow of a survey process executed in a survey processing system according to a first embodiment; [Figure 4] FIG. 4 is an explanatory diagram showing an example of conversion from questionnaire items to question prompts in step ST101 in FIG. 3. [Figure 5] FIG. 4 is an explanatory diagram showing an example of a prompt for evaluating voice in step ST104 in FIG. 3. [Figure 6] FIG. 4 is an explanatory diagram showing an example of an evaluation result regarding the speech recognition result in step ST104 in FIG. 3. [Figure 7] FIG. 4 is an explanatory diagram showing an example of a speech recognition result input from the answer control unit to the text analysis unit in step ST106 in FIG. 3. [Figure 8] FIG. 4 is an explanatory diagram showing an example of a response (questionnaire response) output from the text analysis unit to the response control unit in step ST106 in FIG. 3. [Figure 9] FIG. 4 is an explanatory diagram showing an example of answer information stored in the answer storage unit in step ST108 in FIG. 3. [Figure 10] An explanation of an example of a prompt for answer aggregation input from the answer aggregation unit to the text analysis unit in step ST109 in FIG. [Figure 11] FIG. 4 is an explanatory diagram showing an example of specific opinions and impressions input from the response collection unit to the text analysis unit in step ST109 in FIG. 3. [Figure 12] FIG. 4 is an explanatory diagram showing an example of the results of unifying spelling variations output from the text analysis unit to the answer aggregation unit in step ST109 in FIG. 3. [Figure 13] FIG. 4 is an explanatory diagram showing an example of aggregated answers stored in the aggregated answer storage unit in step ST109 in FIG. 3. [Figure 14] FIG. 4 is an explanatory diagram showing an example of the results of data analysis in step ST110 in FIG. 3. [Figure 15] FIG. 4 is an explanatory diagram showing an example of a response determined by the alert determination unit 55 to require an alert in step ST111 in FIG. 3. [Figure 16] FIG. 4 is an explanatory diagram showing an example of a speech recognition result in which an alert is determined to be necessary by the alert determination unit 55 in step ST111 in FIG. 3. [Figure 17] FIG. 4 is an explanatory diagram showing an example of a prompt for summarization input from the alert summarization control unit to the text analysis unit in step ST111 in FIG. 3. [Figure 18] FIG. 4 is an explanatory diagram showing an example of a response (summary result) output from the text analysis unit to the alert summary control unit in step ST111 in FIG. 3. [Figure 19] FIG. 4 is an explanatory diagram showing an example of an alert notification output from the alert unit to the alert receiving unit in step ST111 in FIG. 3. [Figure 20] FIG. 4 is an explanatory diagram showing another example of the speech recognition result input from the answer control unit to the text analysis unit in step ST106 in FIG. 3. [Figure 21] FIG. 4 is a diagram illustrating another example of a question prompt input from the answer collection unit to the text analysis unit in step ST106 of FIG. [Figure 22] FIG. 4 is an explanatory diagram showing an example of a rule for determining a notification destination based on an alert condition in step ST107 in FIG. 3. [Figure 23] FIG. 4 is an explanatory diagram showing an example of an alert notification output from the alert unit to the alert receiving unit in step ST111 in FIG. 3. [Figure 24] An explanatory diagram showing an example of an answer (explanation of the problem) output from the text analysis unit to the answer control unit. [Figure 25] Overall configuration diagram of a questionnaire processing system according to a second embodiment [Figure 26] FIG. 10 is a flow chart showing the flow of a survey change process performed in the survey processing system according to the second embodiment. [Figure 27] An explanatory diagram showing an example of a comparative answer set [Figure 28] Overall configuration diagram of a questionnaire processing system according to a third embodiment [Figure 29] FIG. 29 is an explanatory diagram showing an example of a correction prompt input from the text correction unit shown in FIG. 28 to the text analysis unit; [Figure 30] FIG. 29 is an explanatory diagram showing an example of a speech recognition result (before correction) input from the text correction unit shown in FIG. 28 to the text analysis unit and a speech recognition result (after correction) output from the text analysis unit to the recognition result correction unit. [Figure 31] Overall configuration diagram of a questionnaire processing system according to a fourth embodiment [Figure 32] Overall configuration diagram of a questionnaire processing system according to a fifth embodiment [Figure 33] FIG. 33 is an explanatory diagram showing an example of a speech recognition result synthesized by the recognition result synthesis unit shown in FIG. 32. [Figure 34] Overall configuration diagram of a questionnaire processing system according to a sixth embodiment [Figure 35] Overall configuration diagram of a questionnaire processing system according to a seventh embodiment [Figure 36] FIG. 20 is an explanatory diagram showing an example of a speech recognition result by a speech recognition unit according to an eighth embodiment. [Figure 37] FIG. 20 is an explanatory diagram showing an example of conversion from a questionnaire item to a question prompt by a prompt conversion unit according to the eighth embodiment. [Figure 38] FIG. 20 is an explanatory diagram showing an example of an answer output from a text analysis unit to an answer control unit according to the eighth embodiment. [Figure 39] FIG. 20 is an explanatory diagram showing an example of a speech recognition result by a speech recognition unit according to a ninth embodiment; [Figure 40] FIG. 20 is an explanatory diagram showing an example of conversion from a questionnaire item to a question prompt by a prompt conversion unit according to the ninth embodiment. [Figure 41] FIG. 20 is an explanatory diagram showing an example of an answer output from a text analysis unit to an answer control unit according to the ninth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] The first invention made to solve the above problem is: Regarding the subject conversation between the customer and the operator, from the perspective of the customer who is a participant in said conversation Survey of answerA questionnaire processing system for acquiring the questionnaire answer the one or more processors execute a process for acquiring questionnaire items as questions related to the questionnaire; Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversation By instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The answer output unit, which is capable of outputting answers to questions posed to voice, is instructed to output answers to the questionnaire items in the voice recognition results, and the answers to the questionnaire items are obtained from the answer output unit.
[0014] This makes it possible to obtain the results of a survey regarding the response of an operator who spoke with a customer using a simple configuration.
[0015] In addition, a second invention is configured such that the one or more processors instruct a speech recognition evaluation unit capable of evaluating a result of speech recognition to evaluate the result of speech recognition, thereby obtaining the evaluation result from the speech recognition evaluation unit, and determining, based on the evaluation result, whether or not to instruct the answer output unit to output the answer.
[0016] This makes it possible to eliminate speech recognition results with low evaluations based on the evaluation results (i.e., not instruct the output of answers for speech recognition results with low evaluations), so that appropriate answers to questionnaire items can be obtained from the answer output unit based only on call voices with high speech recognition accuracy.
[0017] In addition, a third invention is configured such that the questionnaire items are set based on human input operations, the answer output unit is configured using a large-scale language model, and the one or more processors convert the questionnaire items into prompts, and the prompts instruct the answer output unit to output answers to the questionnaire items into the speech recognition results.
[0018] This allows appropriate answers to the questionnaire items to be obtained from the answer output unit based on the prompts into which the questionnaire items have been converted.
[0019] In addition, a fourth invention is configured such that the answers obtained from the answer output unit include the opinions and impressions of the customer, and the one or more processors sequentially accumulate the answers obtained from the answer output unit and instruct the answer output unit to unify the spelling variations in the sentences of the accumulated multiple answers, thereby obtaining multiple answers from the answer output unit that include sentences with unified spelling variations.
[0020] This allows the user of the responses to the questionnaire items (for example, the person managing the operators) to easily understand the customer's impressions of the call with the operator based on the customer's opinions and impressions, which are consistent in the variations in notation in the response text.
[0021] In addition, a fifth invention is configured such that the one or more processors instruct a plurality of speech recognition units capable of evaluating speech recognition results to perform speech recognition on the call voice, respectively, thereby acquiring the speech recognition results from the plurality of speech recognition units, determining whether the plurality of speech recognition results match, and instructing the answer output unit to output the answer only when it is determined that the plurality of speech recognition results match.
[0022] This allows answers to questionnaire items to be obtained from the answer output unit only when there is a high possibility that the voice recognition results are appropriate (i.e., the voice recognition results obtained from multiple voice recognition units match), making it possible to stably obtain appropriate answers to questionnaire items.
[0023] In addition, a sixth invention is configured such that the one or more processors determine whether an alert is necessary based on the answers to the questionnaire items obtained from the answer output unit, and if it is determined that the alert is necessary, output an alert regarding the answers to the questionnaire items to a pre-set alert receiving unit.
[0024] This allows the user of the answers to the questionnaire items (for example, the operator manager using the alert receiving unit) to easily understand answers that require confirmation (for example, answers regarding calls with operators that customers are dissatisfied with) by receiving the alert.
[0025] The seventh invention is the above-mentioned A conversation is an online conversation between a first user and a second user, and the first user may share the conversation with the second user with the first user. The one or more processors execute the process by the terminal, and when the alert is output, the one or more processors execute the process by the terminal. First User Controlling calls to the terminal The end The terminal control unit is configured to instruct the terminal control unit to suspend further calls for a predetermined time.
[0026] This allows the operator who made the call for which the alert was issued to have a break (i.e., calls to the relevant operator terminal are suspended for a predetermined period of time), thereby reducing the psychological anxiety of the operator.
[0027] In addition, an eighth invention is configured such that, when instructing the answer output unit to output an answer to the questionnaire item in the speech recognition result, the one or more processors instruct the answer output unit to output a basis for the answer.
[0028] This allows the user of the answers to the questionnaire items (for example, the person who manages the operators) to easily understand the intentions of the customers regarding the answers based on the basis of the answers.
[0029] In addition, a ninth aspect of the present invention is a method for implementing the one or more processors, the conversation is an online conversation; If hold music is included ,before and removing the hold tone, and transmitting to the voice recognition unit the voice data from which the hold tone data has been removed. conversation The configuration is such that voice recognition is instructed for the voice.
[0030] According to this, since the hold tone included in the call voice is eliminated, it is possible to prevent the hold tone from adversely affecting the answer output from the answer output unit.
[0031] In addition, a tenth aspect of the present invention is a method for implementing the one or more processors, First User Regarding the audio No. 1 Audio and Second User Regarding the audio No. 2 and the speech is converted into text by the speech recognition unit. No. 1 Audio and text versions of the above No. 2 The speech recognition results for instructing the answer output unit to output the answers to the questionnaire items include: First User Audio and front The second user The voice is configured to include the voices in a manner that allows them to be distinguished from each other.
[0032] This makes it possible to obtain appropriate answers to the questionnaire items from the answer output unit based on the call voice data (that is, the voice recognition result) from which the speaker has been distinguished.
[0033] The eleventh invention is: Regarding the subject conversation between the customer and the operator, from the perspective of the customer who is a participant in said conversation Survey of answer a questionnaire processing method for acquiring a questionnaire item, the method comprising: one or more computers acquiring questionnaire items as questions related to the questionnaire; Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversationBy instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The answer output unit, which is capable of outputting answers to questions posed to voice, is instructed to output answers to the questionnaire items in the voice recognition results, and the answers to the questionnaire items are obtained from the answer output unit.
[0034] This makes it possible to obtain the results of a survey regarding the response of an operator who spoke with a customer using a simple configuration.
[0035] The twelfth invention is: Between the customer and the operator Target conversations , from the perspective of a customer who is a participant in the conversation Survey of answer A questionnaire processing program that causes a computer to execute information processing to acquire ,before The information processing includes acquiring questionnaire items as questions regarding the questionnaire, Conversational voice of the subject and transmits the above to a speech recognition unit capable of performing speech recognition. conversation By instructing speech recognition for the speech, the speech recognition unit outputs the speech recognition result converted into text. conversation Audio was captured and converted to text conversation The configuration includes a procedure for acquiring answers to the questionnaire items from an answer output unit capable of outputting answers to questions asked to voice by instructing the answer output unit to output answers to the questionnaire items in the voice recognition results.
[0036] This makes it possible to obtain the results of a survey regarding the response of an operator who spoke with a customer using a simple configuration.
[0037] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0038] (First embodiment) As shown in Figure 1, the survey processing system 1 according to the first embodiment includes an operator terminal 3, a call unit 5, a voice receiving unit 7, a voice recognition unit 9, a survey setting unit 11, a text analysis unit 13, an analysis display unit 15, an alert receiving unit 17, and a main control unit 19.
[0039] The questionnaire processing system 1 is used, for example, to obtain the results of a questionnaire about an operator who answered a telephone inquiry from a customer at a call center. The questionnaire processing system 1 can obtain answers to questions about the questionnaire not from an actual customer, but from a virtual customer that simulates the customer who made the call. The customer can call the call center using the telephone 20 to talk to an operator.
[0040] 1 shows only one operator terminal 3 used by one operator (not shown). However, a call center usually has multiple operators, each of whom is assigned an operator terminal 3. Also, FIG. 1 shows only one telephone set 20 used by one customer (not shown). However, inquiries to a call center are usually made via telephone sets used by multiple customers.
[0041] The operator terminal 3 is configured as a PC, a smartphone, a tablet terminal, etc. However, the operator terminal 3 is not limited to these, and may be configured as any device that has at least a call function.
[0042] 2, the operator terminal 3 includes a display device 21 such as a liquid crystal panel, an input device 22 such as a touch panel, a keyboard, and a mouse, a microphone 23, and a speaker 24. The microphone 23 and the speaker 24 may be headsets. A device having a configuration similar to that of the operator terminal 3 may be used as the telephone 20.
[0043] The operator can use the microphone 23 and speaker 24 of the operator terminal 3 to talk to the customer using the telephone 20. The operator terminal 3 uses an application such as a browser to display various screens that the operator should view on the display device 21 based on display information sent from the call unit 5. The operator terminal 3 can detect screen operations performed by the operator using the input device 22 and send the operation information to the call unit 5.
[0044] The call unit 5 relays the transmission and reception of voice data between the telephone 20 used by the customer and the operator terminal 3 used by the operator, enabling calls between the customer and the operator. The call unit 5 is realized, for example, by a known cloud service that provides cloud contact center functions. Alternatively, the call unit 5 may be realized by a telephone system that can utilize business phone functions over an internet line by building a "PBX (main unit)" function on the cloud.
[0045] The voice receiving unit 7 receives (collects) the voice of the call between the customer and the operator (hereinafter, sometimes simply referred to as "call voice") made via the call unit 5. The voice receiving unit 7 is realized, for example, by a known cloud service that provides a call recording function. Note that the voice receiving unit 7 may also be realized as part of a known cloud contact center function.
[0046] The voice recognition unit 9 executes a process of converting the call voice acquired by the voice receiving unit 7 into text by voice recognition (hereinafter referred to as "voice recognition process") in response to an instruction from the main control unit 19. The voice recognition unit 9 is realized, for example, by a known cloud service that provides a function of converting voice into text.
[0047] The survey setting unit 11 is configured with a PC, smartphone, tablet terminal, or the like. The survey setting unit 11 is used by a survey person who belongs to, for example, a customer management department. By accessing the main control unit 19 from the survey setting unit 11, the survey person can set questions for a survey regarding the operator who spoke with the customer. The set questions are sent to the main control unit 19.
[0048] The text analysis unit 13 is realized by a known cloud service that uses large language models (LLMs), such as Google Gemini (registered trademark), Microsoft Bing AI (registered trademark), or ChatGPT. The text analysis unit 13 can execute various processes in response to instructions from the main control unit 19. The multiple processes executed by the text analysis unit 13 may each be implemented by multiple different cloud services. Note that the processes may also be implemented by implementing a large language model on an on-premise server, rather than by a cloud service.
[0049] For example, the text analysis unit 13 (an example of a voice recognition evaluation unit) can execute a process (hereinafter referred to as "voice recognition evaluation process") to output an evaluation of the voice recognition result of the voice recognition unit 9 (i.e., the call voice converted into text by the voice recognition process).
[0050] Furthermore, the text analysis unit 13 (an example of an answer output unit) can execute a process (hereinafter referred to as an "answer output process") of outputting answers to questions prepared in advance based on the speech recognition results of the speech recognition unit 9. The prepared questions include questions related to a questionnaire set by the questionnaire setting unit 11 (hereinafter referred to as "questionnaire items" as necessary). The purpose of the questionnaire is to have customers (more precisely, virtual customers realized by the text analysis unit 13) answer questions about, for example, their satisfaction with or dissatisfaction with a call with an operator. Furthermore, the text analysis unit 13 can execute a process (hereinafter referred to as an "answer aggregation process") of standardizing variations in text notation for multiple answers to questionnaire items.
[0051] In addition, as will be described later, when it becomes necessary to issue an alert to the operator's manager (e.g., superior) regarding the answers obtained to the questionnaire items, the text analysis unit 13 can execute a process of generating a summary of the corresponding call voice (hereinafter referred to as "alert summary generation process").
[0052] The analysis display unit 15 is composed of a PC, smartphone, tablet terminal, or the like. The analysis display unit 15 is used by an analyst, such as an operator manager, who is in charge of analyzing the survey results. In the analysis display unit 15, the analyst can instruct the main control unit 19 regarding data analysis of the accumulated responses. Such data analysis includes, for example, statistical processing of the accumulated responses and the creation of tables and graphs. The analyst can also instruct the main control unit 19 to display the accumulated responses and the results of the data analysis.
[0053] The alert receiving unit 17 is configured with a PC, smartphone, tablet terminal, or the like. The alert receiving unit 17 is used by a trouble-shooting person such as an operator manager. The alert receiving unit 17 allows the trouble-shooting person to receive alerts output from the main control unit 19. This allows the trouble-shooting person to respond to customer complaints regarding calls with operators and provide support to the operators based on the content of the alert.
[0054] The main control unit 19 includes a questionnaire input unit 31, a prompt conversion unit 33, a voice input unit 35, a voice recognition control unit 37, a voice recognition evaluation control unit 39, an answer target selection unit 41, an answer control unit 43, an answer storage unit 45, an answer aggregation unit 47, an aggregated answer storage unit 49, an analysis instruction input unit 51, a data analysis unit 53, an alert determination unit 55, an alert summary control unit 57, and an alert unit 59. The main control unit 19 is configured to be able to communicate with the voice receiving unit 7, the voice recognition unit 9, the text analysis unit 13, the questionnaire setting unit 11, the analysis display unit 15, the alert receiving unit 17, etc. via a known communication network.
[0055] The questionnaire input unit 31 accepts input (an example of a human input operation) related to the questionnaire from the questionnaire setting unit 11. The questionnaire items to be input can be changed as appropriate by the person in charge of the questionnaire.
[0056] A prompt must be input to instruct the large-scale language model. However, survey personnel generally do not know what to write as a prompt. Therefore, the prompt conversion unit 33 converts the survey items input from the survey setting unit 11 into prompts (hereinafter referred to as "question prompts" as necessary). The question prompts include character strings that clearly convey specific tasks and requirements to the text analysis unit 13. The converted prompts are also input to the answer control unit 43. The converted prompts may include pairs of provisional speech recognition results and example answers to realize few-shot learning.
[0057] The voice input unit 35 receives input of call voice (voice data) from the voice receiving unit 7. For example, when one call between a customer and an operator (i.e., one call communication between the telephone 20 and the operator terminal 3) is completed, the call voice is input to the voice input unit 35 from the voice receiving unit 7. Note that the call voice may be input sequentially from the voice receiving unit 7 in order to increase the processing speed of voice recognition.
[0058] The voice recognition control unit 37 acquires the call voice input to the voice input unit 35, and instructs the voice recognition unit 9 to execute voice recognition processing on the call voice. As a result, the voice recognition control unit 37 acquires the call voice converted into text (i.e., text data of the call voice) as the result of the voice recognition processing by the voice recognition unit 9 (hereinafter, sometimes simply referred to as the "voice recognition result"). Note that, in order to increase the processing speed of voice recognition, the voice recognition unit 9 may be instructed to execute voice recognition processing on call voices input sequentially, and the voice recognition results may be combined for each call.
[0059] The voice recognition control unit 37 can obtain information about noisy lines, such as calls from mobile phone numbers, from the call unit 5, and exclude the voice of calls using that line from the voice recognition processing. The voice recognition control unit 37 can also obtain information about wrong numbers, such as calls from phone numbers not registered in advance with the customer, from the call unit 5, and exclude the voice of calls related to those wrong numbers from the voice recognition processing.
[0060] The voice recognition evaluation control unit 39 instructs the text analysis unit 13 to execute a voice recognition evaluation process for the voice recognition result obtained by the voice recognition control unit 37. As a result, the voice recognition evaluation control unit 39 acquires an evaluation result for the voice recognition result by the voice recognition unit 9 (i.e., the result of the voice recognition evaluation process) from the text analysis unit 13. The evaluation result serves as an index of the quality of the text of the call voice obtained by the voice recognition process (here, whether it is suitable as a target for answer output process by the text analysis unit 13).
[0061] Based on the evaluation result of the speech recognition result obtained by the speech recognition evaluation control unit 39, the answer target selection unit 41 determines whether or not to instruct the text analysis unit 13 to execute an answer output process (i.e., output an answer) for the speech recognition result to be evaluated (i.e., the text of the call voice). If the content of the call voice is not accurately reflected in the speech recognition result (i.e., the evaluation of the speech recognition result is low), this may have a negative impact on the results of the answer output process and ultimately the results of the data analysis. Therefore, the answer target selection unit 41 excludes speech recognition results with low evaluations from the targets of the answer output process (i.e., determines that the text analysis unit 13 should not be instructed to execute the answer output process).
[0062] The answer control unit 43 acquires a speech recognition result determined by the answer target selection unit 41 to be instructed to instruct the text analysis unit 13 to execute an answer output process. The answer control unit 43 also acquires a prompt corresponding to the speech recognition result from the prompt conversion unit 33. Then, the answer control unit 43 instructs the text analysis unit 13 to execute an answer output process (i.e., output answers to the questionnaire items) using the acquired speech recognition result and the corresponding prompt. As a result, the answer control unit 43 acquires answers to the questionnaire items regarding the speech recognition result from the text analysis unit 13. Note that answers to the questionnaire items may be output in a specific language (e.g., Japanese) by instructing the text analysis unit 13 to output them in a specific language, regardless of the language used in the conversation between the customer and the operator or the language used in the prompt.
[0063] The answer information acquired by the answer control unit 43 is sequentially stored in the answer accumulation unit 45. The answer accumulation unit 45 includes a database constructed on the cloud, for example.
[0064] The answer aggregation unit 47 instructs the text analysis unit 13 to execute an answer aggregation process for the answer information accumulated in the answer accumulation unit 45. As a result, the answer aggregation unit 47 acquires from the text analysis unit 13 a plurality of answers for which the answer aggregation process has been performed (hereinafter referred to as "aggregated answers").
[0065] The aggregated answers include multiple answers in which variations in the text notation of free-form answers such as customer opinions and impressions are unified. For example, the answer aggregation unit 47 can instruct the text analysis unit 13 to execute the answer aggregation process when the main control unit 19 (here, the analysis instruction input unit 51) receives an instruction regarding data analysis from the analysis display unit 15. Alternatively, the answer aggregation unit 47 may instruct the text analysis unit 13 to execute the answer aggregation process when a predetermined amount of answer information has been accumulated in the answer storage unit 45. Note that the answer aggregation unit 47 may also instruct the text analysis unit 13 to execute the answer aggregation process at a specific scheduled time, such as midnight on the first day of each month.
[0066] Information on the aggregated answers acquired by the answer aggregation unit 47 is sequentially stored in the aggregated answer storage unit 49. The aggregated answer storage unit 49 includes a database constructed on the cloud, for example.
[0067] The analysis instruction input unit 51 receives instructions regarding data analysis from the analysis display unit 15. The input instructions regarding data analysis can be changed as appropriate by the analyst.
[0068] The data analysis unit 53 performs data analysis on the aggregated responses stored in the aggregated response storage unit 49 based on instructions related to data analysis input to the analysis instruction input unit 51. As part of the data analysis, the data analysis unit 53 can process data related to the aggregated responses using statistical methods and display the results as tables or graphs using visualization methods. The data analysis unit 53 also displays the results of such data analysis on the analysis display unit 15. For example, the data analysis unit 53 can perform data analysis using a cloud-based BI (Business Intelligence) tool and display the results on the analysis display unit 15.
[0069] The alert determination unit 55 determines whether an alert is necessary based on the answers to the questionnaire items acquired by the answer control unit 43. For example, the alert determination unit 55 can determine that an alert is necessary when the answers to the questionnaire items include information indicating customer dissatisfaction.
[0070] The alert summary control unit 57 instructs the text analysis unit 13 to execute an alert summary generation process for a response determined by the alert determination unit 55 to require an alert. This allows the alert summary control unit 57 to acquire summary result information (hereinafter referred to as "alert information" as necessary) related to the corresponding call voice from the text analysis unit 13. The alert information includes, for example, the name of the store used by the customer (or the name of the product purchased by the customer or the service received by the customer), customer identification information (e.g., the customer's name), operator identification information (e.g., the operator's name), the content of the customer's inquiry, and the content of the operator's response. If the call unit 5 is configured using a cloud service PBX, the customer identification information and the operator identification information may be managed by the cloud PBX. In this case, the customer and operator identification information may be acquired directly from the voice receiving unit 7.
[0071] When the alert determination unit 55 determines that an alert is necessary, the alert unit 59 outputs (or transmits) the alert information acquired by the alert summary control unit 57 to the alert receiving unit 17. This allows a troubleshooter using the alert receiving unit 17 to check the contents of the alert. However, the alert unit 59 may output a response for which it has been determined that an alert is necessary to the alert receiving unit 17 without adding a summary. In that case, the alert summary control unit 57 may be omitted from the main control unit 19.
[0072] At least some of the functions of each unit of the main control unit 19 as described above can be realized, for example, by cloud computing (i.e., servers, storage, network infrastructure, databases, software, etc. in a cloud environment). The main functions of the main control unit 19 can be realized by one or more processors provided in one or more servers executing predetermined control programs. Furthermore, at least some of the functions of each unit of the main control unit 19 may be realized by edge computing (edge servers, etc.).
[0073] Next, the flow of the survey processing executed in the survey processing system 1 according to the first embodiment will be described with reference to Fig. 3 etc. Here, the survey processing is processing for obtaining the results of a survey regarding an operator who spoke with a customer (including the customer's answers and the results of data analysis of the answers).
[0074] In the questionnaire processing system 1, first, when questionnaire items are input from the questionnaire setting unit 11 to the questionnaire input unit 31, the prompt conversion unit 33 converts the questionnaire items into question prompts (ST101).
[0075] In step ST101, as shown in FIG. 4, for example, the questionnaire items (before conversion) set by the person in charge of the questionnaire are converted into question prompts (after conversion).
[0076] The question prompt includes information such as the respondent's position (here, a hypothetical customer) and the type of text being asked (here, a call center transcript), etc. The question prompt also includes questions regarding, for example, the operator's response, the operator's language, the operator's understanding, the resolution of the customer's problems and questions, re-use (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction with the operator's response, etc. (an example of the basis for the answer), and the customer's opinions and impressions.
[0077] For example, as shown in Figure 4, the prompt begins with the following text: "You are a customer who called the call center. What you are about to enter is a transcript from the call center. Please answer the survey questions from the customer's perspective." This allows the user to virtually answer the survey from the perspective of the customer who made the call.
[0078] Next, add "Please output in JSON format using the following key" to the prompt. This will output the answer in a computer-friendly JSON format.
[0079] Next, add "# Key, Question" to the prompt. This is an instruction to enter the key phrase in JSON format and the question content.
[0080] Furthermore, the key phrases and question contents are extracted as pairs from the questionnaire items before conversion and written in the prompt. For example, Q1.Response How was the call center response? Choose from "Satisfied," "Almost Satisfied," "Average," "Slightly Dissatisfied," and "Dissatisfied." Please answer. For the above questionnaire item, the string after QX. is extracted as the Key, and the string after the line break after the Key is recognized as the question, converted as shown below, and added to the prompt. Response, How was the call center's response? Please choose from "Satisfied," "Mostly Satisfied," "Average," "Slightly Dissatisfied," or "Dissatisfied."
[0081] The question prompt may include, for example, attribute information (e.g., age, gender) of the actual customer contained in the call audio. If the call unit 5 is configured using a cloud service PBX, the customer's identification information may be managed by the cloud PBX, and the prompt conversion unit 33 may be able to acquire the identification information. For example, if the caller is a "male under the age of 10," the prompt may be converted to match the customer's identification information, such as "You are a male under the age of 10 who called the call center. What you are about to enter is a transcript from the call center. Please answer the survey questions from your perspective." In this way, by instructing the text analysis unit to have a role that includes the caller's own attribute information, it is possible to obtain survey responses that more closely resemble those of the caller. The question prompt may be adjusted so that, even when similar questions are asked of an actual customer who made the call audio, responses similar to those of a virtual customer (i.e., the text analysis unit 13) are obtained. The question prompt may also be configured to prompt a specific score for each question.
[0082] It should be noted that once the processing of step ST101 is performed, it can be omitted (that is, the same question prompt is used repeatedly) until the questionnaire item is changed.
[0083] When one call between the customer and the operator ends (Yes in ST102), the call voice is input from the voice receiving unit 7 to the voice input unit 35, and the voice recognition control unit 37 obtains the voice recognition result of the call voice from the voice recognition unit 9 (ST103).
[0084] Next, the voice recognition evaluation control unit 39 acquires the evaluation result regarding the voice recognition result acquired in step ST103 from the text analysis unit 13 (ST104).
[0085] In step ST104, a prompt such as that shown in FIG. 5 (hereinafter referred to as a "speech evaluation prompt") is input from the speech recognition evaluation control unit 39 to the text analysis unit 13 together with the speech recognition result.
[0086] The prompt for voice evaluation includes information such as the respondent's position (here, a hypothetical professional editor) and the type of text to be questioned (here, a transcript of a conversation generated by speech recognition). The prompt for questioning also includes questions regarding, for example, the accuracy of the speech recognition, the reasons for the professional editor's dissatisfaction with the accuracy of the speech recognition, and the professional editor's opinions and impressions.
[0087] 6, for example, the evaluation result regarding the speech recognition result is output from the text analysis unit 13 to the speech recognition evaluation control unit 39. The evaluation result includes answers to each question in the speech evaluation prompt.
[0088] Next, the answer target selection unit 41 determines whether the speech recognition result is appropriate as an answer target based on the evaluation result obtained by the speech recognition evaluation control unit 39, and excludes inappropriate speech recognition results from targets for answer output processing (ST105). For example, as shown in Fig. 6, speech recognition results for which the answer item "speech recognition accuracy" is "slightly unsatisfactory" or "unsatisfactory" are determined to be inappropriate and excluded. In this way, only highly evaluated speech recognition results are selected as targets for answer output processing.
[0089] Thereafter, the answer control unit 43 obtains answers to the questionnaire items regarding the voice recognition results from the text analysis unit 13 using the voice recognition results and the corresponding question prompts that have been determined to be appropriate by the answer target selection unit 41 (ST106).
[0090] In step ST106, the speech recognition result (i.e., the text to be answered) and the question prompt determined to be appropriate, for example, as shown in FIG. 7, are input from the answer control unit 43 to the text analysis unit 13, and the text analysis unit 13 is instructed to execute an answer output process. This speech recognition result is acquired in step ST103, and is determined to be appropriate as a subject of an answer in ST105 based on the evaluation result in step ST104. Note that the question prompt corresponding to this speech recognition result is the same as the question prompt (after conversion) shown in FIG. 4.
[0091] 8, for example, are output from the text analysis unit 13 to the answer control unit 43. The answers that are output correspond to the respective questions in the question prompt.
[0092] Next, the alert determination unit 55 determines whether an alert is necessary based on the response obtained in step ST106 (ST107). If it is determined that an alert is unnecessary (No in ST107), the response control unit 43 stores the information of the obtained response in the response accumulation unit 45 (ST108).
[0093] In step ST108, information on a plurality of answers is stored in the answer storage unit 45, as shown in Fig. 9, for example. The answer information includes identification information (here, a contact ID) for identifying the call between the customer and the operator, as well as answers to each question associated therewith (here, the operator's response, the operator's wording, the operator's understanding, whether to use the service again, the reason for the customer's dissatisfaction with the operator's response, etc., and the customer's opinions and impressions, etc.).
[0094] Thereafter, the response aggregation unit 47 acquires aggregated responses for the information of the responses accumulated in the response accumulation unit 45 (ST109). The acquired information of the aggregated responses is saved (accumulated) in the aggregated response accumulation unit 49.
[0095] In step ST109, for example, a prompt for answer collection shown in FIG. 10 is input from the answer control unit 43 to the text analysis unit 13, and the text analysis unit 13 is instructed to execute the answer collection process.
[0096] The response aggregation prompt includes information such as the respondent's position (here, a hypothetical professional editor), the type of text being questioned (here, customer opinions and impressions collected through a questionnaire), and specific instructions (aggregating similar opinions from the questionnaire results).The response aggregation prompt also includes instructions regarding the input format (here, specific opinions and impressions) and the output format (here, an aggregation of similar opinions and impressions).
[0097] In addition, in step ST109, together with the above-mentioned response aggregation prompt, part of the response information stored in the response storage unit 45 (free-form response items among the questionnaire items), for example, the contact IDs shown in Figure 11 and specific opinions and impressions from multiple customers (text to be standardized for spelling variations) are input from the response control unit 43 to the text analysis unit 13.
[0098] In step ST109, the result of standardizing the spelling variations in the sentences (here, similar sentences relating to opinions and impressions are aggregated) such as that shown in Fig. 12 is output from the text analysis unit 13 to the response aggregation unit 47. In this result of standardizing the spelling variations, information on the aggregated opinions and impressions (i.e., the spelling variations in the text are standardized) is shown in association with the identification information of the corresponding multiple calls (here, contact IDs).
[0099] In step ST109, as shown in FIG. 13, for example, information on multiple aggregated answers is stored in the aggregated answer storage unit 49, combining the results of standardizing spelling variations output from the text analysis unit 13 to the answer aggregation unit 47 with the answer information accumulated in the answer storage unit 45 based on the identification information of multiple calls (contact IDs in this case). The aggregated answers include opinions and impressions after aggregation based on the results of standardizing spelling variations, in addition to the content of the answer information shown in FIG. 9. This aggregation process makes it easier for the data analysis unit 53 to perform statistical processing.
[0100] It should be noted that step ST109 does not always need to be executed when answers based on the answer aggregation process are acquired from the text analysis unit 13 (ST106), and may be omitted in some cases. In particular, if the questionnaire does not include free-form questions such as specific opinions and impressions from customers, it is effective to omit the aggregation process because there is no need to correct spelling variations.
[0101] Thereafter, the data analysis unit 53 performs data analysis such as statistical processing on the aggregated response data based on instructions for data analysis from the analyst input through the analysis instruction input unit 51 from the analysis display unit 15 (ST110). For example, if an instruction is given to create a pie chart showing the number of opinions and impressions after aggregation for the aggregated responses in FIG. 13, the number of opinions and impressions after aggregation is counted and tallied to create a pie chart, which is then displayed on the analysis display unit 15, as shown in FIG. 14. The data analysis in step ST110 can be started when instructions for data analysis are input to the analysis instruction input unit 51.
[0102] On the other hand, if it is determined in the above-mentioned step ST107 that an alert is necessary (Yes), the alert section 59 outputs (or transmits) alert information to the alert receiving section 17 (ST111).
[0103] The answers of the answer control unit 43 that are determined to require an alert in step ST107 include, for example, as shown in FIG. 15, answers in which the questionnaire result for the item on the call center's response matches "dissatisfied."
[0104] In step ST111, the voice recognition result resulting in the input of a questionnaire response that has been determined to require an alert, as shown in FIG. 16, and a summary prompt, as shown in FIG. 17, are input from the alert summarization control unit 57 to the text analysis unit 13.
[0105] Also, in step ST111, a response (summarization result) in response to the summarization prompt is output from the text analysis unit 13 to the alert summarization control unit 57, as shown in FIG.
[0106] Furthermore, in step ST111, as shown in Fig. 19, for example, an alert notification including the contact ID, call date and time, and operator name acquired from the call unit, the summary result output from the text analysis unit 13 to the alert summary control unit 57, and the survey results is output from the alert unit 59 to the alert receiving unit 17. As a result, the person in charge of troubleshooting is notified immediately after the call about any telephone response that the customer is dissatisfied with. This allows the person in charge of troubleshooting to immediately follow up with the customer, preventing the problem from escalating into a serious problem.
[0107] Here, the alert notification includes information such as the date and time of the call, call identification information (here, contact ID), operator identification information (e.g., the operator's name), customer identification information (e.g., the name and name of the store the customer belongs to), the operator's response, the customer's evaluation of the operator's response, the customer's evaluation of the operator's language, the customer's evaluation of the operator's level of understanding, the customer's evaluation of the resolution of the problem, repeat use (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction, and the customer's opinions and impressions.
[0108] The notification destination of the alert may be changed depending on the content of the survey responses. If the customer is dissatisfied with the service, the troubleshooter can be notified, allowing for immediate follow-up with the customer. If the customer is angry, the operator's manager can be notified as well as the troubleshooter, allowing the operator manager to immediately follow up with the operator who handled the call and alleviate the operator's psychological anxiety. Furthermore, if the customer is grateful, the operator who handled the call can be notified, thereby improving the operator's motivation.
[0109] Furthermore, if the person receiving the alert can easily check the relevant part of the audio that caused the problem when an alert is issued, they will be able to take more appropriate action, such as following up with the customer. Angry calls, in particular, can last more than 30 minutes, making it difficult to listen to the entire audio again in order to follow up. By being able to easily play back the audio of the problematic part when an alert is issued, the person receiving the notification can easily understand the problematic part of the call that caused the alert.
[0110] In step ST103, the voice recognition control unit 37 acquires the voice recognition result of the call voice input to the voice input unit 35 from the voice recognition unit 9. The voice recognition result includes the duration of each utterance by the customer and the operator, for example, as shown in FIG.
[0111] In step ST106, the response control unit 43 inputs the speech recognition result determined to be appropriate as shown in Fig. 20 (including the times of each utterance by the customer and the operator) and a question prompt as shown in Fig. 21, for example, to the text analysis unit 13. The question prompt includes, for example, questions about the operator's response, the operator's language, the operator's understanding, the resolution of the problem or question that the customer has, the likelihood of using the service again (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction with the operator's response, etc. (an example of the basis for the answer), gratitude to the operator, the reason for the gratitude, anger towards the operator, the reason for the anger, questions for identifying the problem areas of the anger alert notification (the start and end times of the relevant voices of the customer and the operator, the timing of the customer's anger, etc.), and questions about the customer's opinions and impressions, etc.
[0112] In step ST106, in response to this question prompt, an answer (including an explanation of the problematic part) such as that shown in FIG.
[0113] In step ST107, the alert determination unit 55 determines whether an alert is necessary and the notification destination of the alert based on the answer acquired in step ST106 and the rules (including the alert conditions and the notification destination) as shown in Fig. 22. For example, if the questionnaire answer item "anger" matches "yes," it determines that an alert is necessary and the notification destinations are "aaa@aaa.aaa.jp" and "bbb@bbb.bbb.jp."
[0114] In step ST111, as shown in FIG. 23, for example, an alert notification including the contact ID, call date and time, and operator name obtained from the call unit, the summary results output from the text analysis unit 13 to the alert summary control unit 57, and the survey results is output to the alert receiving unit 17 of the notification destination determined by the alert unit 59.
[0115] Here, the alert notification includes information such as the date and time of the call, call identification information (here, contact ID), operator identification information (for example, the operator's name), customer identification information (for example, the name and name of the store the customer belongs to), the operator's response, the customer's evaluation of the operator's response, the customer's evaluation of the operator's use of language, the customer's evaluation of the operator's level of understanding, the customer's evaluation of the resolution of the problem, likelihood of repeat use (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction, whether the customer is grateful for the operator's response, the reason for that gratitude, whether the customer is angry at the operator's response, the reason for that anger, the timing of that anger, and the customer's opinions and impressions.
[0116] The alert notification also includes a problem section playback button 60. The troubleshooter can press the problem section playback button 60 in the alert notification displayed on the alert receiving unit 17 to play back the problem section in the target call.
[0117] At this time, when the playback button 60 for the problematic portion is pressed, the alert receiving unit 17 plays back the corresponding anger voice from the start time to the end time input to the voice input unit 35. For example, in the case of the answer shown in FIG. 24, the portion from the start time: 00:00:31,000 to the end time: 00:01:03,000 is played back.
[0118] (Second embodiment) Next, a questionnaire processing system 1 according to a second embodiment will be described with reference to FIG. 25 and other figures. Unlike questionnaires targeted at people, this system allows questionnaires to be conducted any number of times by changing the questionnaire items. Therefore, a more effective questionnaire can be created by having the surveyor visually check the results of the questionnaire before and after the change while revising the questionnaire items. In the drawings and explanations of the second embodiment, components similar to those in the first embodiment described above are assigned the same reference numerals as those used in the first embodiment. Furthermore, the questionnaire processing system 1 according to the second embodiment is the same as that of the first embodiment, except for matters specifically mentioned below.
[0119] In the questionnaire processing system 1 according to the second embodiment, it is possible to execute a process (hereinafter referred to as "questionnaire correction process") to correct the questionnaire items set by the person in charge of the questionnaire to more appropriate content based on the answers output from the text analysis unit 13.
[0120] The survey processing system 1 includes a voice storage unit 63 that sequentially stores the voice of each call relayed by the call unit 5. The voice data of each call stored in the voice storage unit 63 is input to the voice input unit 35 as needed.
[0121] The main control unit 19 further includes a comparison information storage unit 65 , a comparison result display unit 67 , and a questionnaire determination unit 69 .
[0122] The aggregated answers stored in aggregated answer storage unit 49 include sets of answers for comparison (hereinafter referred to as "comparison answer sets") obtained by question prompts based on different questionnaires (including different questionnaire items) before and after correction. Comparison information storage 65 saves (stores) the comparison answer sets extracted from aggregated answer storage unit 49.
[0123] The comparison result display unit 67 displays the comparative answer sets stored in the comparison information storage unit 6 on the questionnaire setting unit 11 .
[0124] In the questionnaire setting unit 11, the person in charge of the questionnaire who has checked the comparative answer set can modify the questionnaire items as necessary (that is, select the modified questionnaire items).
[0125] The information on the amendment of the questionnaire item is input from the questionnaire setting unit 11 to the questionnaire determination unit 69. The questionnaire determination unit 69 can input the amended questionnaire item to the questionnaire input unit 31 based on the information on the amendment of the questionnaire item.
[0126] As a result, the questionnaire processing system 1 can cause the text analysis unit 13 to execute the answer output process based on the questionnaire items (question prompts) that have been modified to have more appropriate content.
[0127] Next, the flow of the questionnaire correction process performed in the questionnaire processing system 1 according to the second embodiment will be described with reference to FIG.
[0128] In the questionnaire correction process, ST201 corresponding to step ST101 shown in Fig. 3 is executed. In the following step ST202, a voice file corresponding to each call stored in the voice storage unit 63 is acquired, and the voice file is input to the voice input unit 35. The voice recognition control unit 37 acquires a voice recognition result based on the voice file from the voice recognition unit 9 (ST203).
[0129] Subsequently, steps ST203-ST205 and ST207-ST209 corresponding to steps ST103-ST106 and ST108-ST110 shown in FIG. 3, respectively, are executed.
[0130] Thereafter, when the questionnaire item is corrected by the person in charge of the questionnaire, the corrected questionnaire item is input from the questionnaire setting section 11 to the questionnaire input section 31 (ST210).
[0131] Therefore, main control unit 19 executes processing based on the revised questionnaire (ST211). The processing in step ST211 is the same as steps ST101, ST103-ST106, and ST108-ST110 shown in Fig. 3. Furthermore, in step ST211, the comparison answer set extracted from aggregated answer storage unit 49 is stored in comparison information storage 65.
[0132] Next, the comparison result display unit 67 displays the comparison answer set stored in the comparison information storage 65 on the questionnaire setting unit 11 (ST212). The displayed comparison answer set includes the questionnaire items before and after the change and the corresponding answers (survey results), as shown in Fig. 27, for example.
[0133] Next, when the person in charge of the survey selects the revised survey items (i.e., judges that the revised survey items are more appropriate), the survey item revision information is input from the survey setting unit 11 to the survey determination unit 69 (ST213). As a result, the content of the new survey (survey items) is determined.
[0134] Thereafter, the questionnaire processing system 1 according to the second embodiment can execute processing similar to the questionnaire processing shown in FIG. 3 described above, based on the content of the new questionnaire that has been decided.
[0135] (Third embodiment) Next, a survey processing system 1 according to a third embodiment will be described with reference to FIG. 28 and other figures. In the first embodiment, speech recognition results with poor speech recognition accuracy were excluded from the survey targets. However, there are cases where the number of calls is low and it is not desirable to reduce the number of survey targets. Therefore, by correcting the speech recognition results using a text analysis unit, the accuracy of speech recognition can be improved, and the number of survey targets can be increased while maintaining the quality of the survey responses. In the drawings and explanations related to the third embodiment, components similar to those in the first or second embodiment described above are assigned the same reference numerals as those used in the first or second embodiment. Furthermore, the survey processing system 1 according to the third embodiment is the same as that of the first or second embodiment, except for matters specifically mentioned below.
[0136] In the questionnaire processing system 1 according to the third embodiment, the main control unit 19 further includes a text correction unit 71.
[0137] Here, the text analysis unit 13 can execute a process of proofreading sentences (hereinafter referred to as a "correction process") for the speech recognition result of the speech recognition unit 9. The text correction unit 71 instructs the text analysis unit 13 to execute a correction process for the speech recognition result obtained by the speech recognition control unit 37. As a result, the text correction unit 71 obtains the corrected speech recognition result from the text analysis unit 13.
[0138] When causing the text correcting unit 71 to execute the correction process, the text correcting unit 71 can input a prompt (hereinafter referred to as a "correction prompt") as shown in FIG.
[0139] In addition, the text analysis unit 13 can output a corrected speech recognition result (after correction) to the text correction unit 71 by performing a correction process on the speech recognition result (before correction), for example, as shown in FIG. 30 .
[0140] The voice recognition evaluation control unit 39 can instruct the text analysis unit 13 to execute a voice recognition evaluation process for the voice recognition result after correction acquired by the text correction unit 71 .
[0141] (Fourth embodiment) Next, a survey processing system 1 according to a fourth embodiment will be described with reference to FIG. 31. In a configuration in which an operator can put a call on hold, if the call voice includes a hold tone, the hold tone will reduce the accuracy of speech recognition during the hold period and before and after the hold period. Therefore, by removing the hold tone, the accuracy of speech recognition can be improved, thereby improving the quality of survey responses. In the drawings and explanations of the fourth embodiment, components similar to those in any of the first to third embodiments described above are assigned the same reference numerals as those used in any of the first to third embodiments. Furthermore, the survey processing system 1 according to the fourth embodiment is the same as that in any of the first to third embodiments, except for matters specifically mentioned below.
[0142] In the questionnaire processing system 1 according to the fourth embodiment, the main control unit 19 further includes a hold sound removal unit 74.
[0143] The voice receiving unit 7 receives (collects) the voice of the call between the customer and the operator made by the call unit 5, as well as the start and end of the operator's hold.
[0144] The hold sound removal unit 74 is provided between the voice receiving unit 7 and the voice input unit 35, and receives input of the call voice from the voice receiving unit 7 and the start and end of the operator's hold. If the call voice includes a hold sound for the telephone, the hold sound removal unit 74 uses the operator's hold start and end information to remove data related to the hold sound from the call voice data. As a result, the voice input unit 35 receives the call voice from which the hold sound has been removed from the hold sound removal unit 74.
[0145] With the above configuration, the questionnaire processing system 1 according to the fourth embodiment eliminates the hold tone included in the call voice, and therefore, it is possible to prevent the hold tone from adversely affecting the answers output from the text analysis unit 13.
[0146] (Fifth embodiment) Next, a questionnaire processing system 1 according to a fifth embodiment will be described with reference to FIG. 32 and other figures. When a text analysis unit analyzes a call voice text, it may misidentify the speaker. For example, when an operator thanks a customer, it may mistakenly recognize that the customer is thanking the operator. Misidentification can be prevented by inputting speech recognition results including speaker information to the text analysis unit. In the drawings and explanations related to the fifth embodiment, components similar to those in any of the first to fourth embodiments described above are assigned the same reference numerals as those used in any of the first to fourth embodiments. Furthermore, the questionnaire processing system 1 according to the fifth embodiment is the same as that in any of the first to fourth embodiments, except for matters specifically mentioned below.
[0147] In the questionnaire processing system 1 according to the fifth embodiment, the main control unit 19 further includes a recognition result synthesis unit 75.
[0148] In the survey processing system 1, data on the conversation voice between the customer and the operator acquired by the voice receiving unit 7 (including customer voice data related to the customer's voice and operator voice data related to the operator's voice) is separated into L channel and R channel, respectively, and input to the voice input unit 35.
[0149] The voice recognition control unit 37 instructs the voice recognition unit 9 to execute voice recognition processing on the L channel and R channel voices input to the voice input unit 35 in a separated state. As a result, the voice recognition control unit 37 obtains voice recognition results for the L channel and R channel voices, respectively.
[0150] The recognition result synthesis unit 75 synthesizes the speech recognition results for the L channel and R channel speech acquired by the speech recognition control unit 37. In this case, the synthesized speech recognition result may include the customer's speech and the operator's speech in a distinguishable manner, as shown in Fig. 33, for example.
[0151] With the above configuration, the survey processing system 1 according to the fifth embodiment can obtain appropriate answers to survey items from the text analysis unit 13 based on data of the call voice (i.e., the voice recognition results) in which the speaker (here, the customer and the operator) is distinguished.
[0152] (Sixth embodiment) Next, a questionnaire processing system 1 according to a sixth embodiment will be described with reference to Fig. 34. In the drawings and description of the sixth embodiment, components similar to those in any of the first to fifth embodiments described above are assigned the same reference numerals as those used in any of the first to fifth embodiments. Furthermore, the questionnaire processing system 1 according to the sixth embodiment is the same as that in any of the first to fifth embodiments, except for matters specifically mentioned below.
[0153] The questionnaire processing system 1 according to the sixth embodiment includes, in addition to a voice recognition unit 9A corresponding to the voice recognition unit 9 shown in Fig. 1, a voice recognition unit 9B employing a voice recognition model different from that of the voice recognition unit 9A (for example, a machine learning model different from that of 9A). Note that the questionnaire processing system 1 may include three or more voice recognition units.
[0154] The voice recognition control unit 37 instructs a plurality of voice recognition units (here, the voice recognition unit 9A and the voice recognition unit 9B) to perform voice recognition processing on the voice of the call between the customer and the operator. As a result, the voice recognition control unit 37 can obtain the voice recognition results from the voice recognition unit 9A and the voice recognition unit 9B.
[0155] The voice recognition evaluation control unit 39 can compare the voice recognition results obtained from the voice recognition unit 9A and the voice recognition unit 9B. For example, the voice recognition evaluation control unit 39 can compare multiple (here, two) voice recognition results using WER (Word Error Rate) and evaluate the difference (i.e., determine whether the multiple voice recognition results match).
[0156] The answer target selection unit 41 determines whether or not to instruct the text analysis unit 13 to execute an answer output process for the speech recognition result to be evaluated, based on the evaluation result of the difference between the speech recognition results by the speech recognition evaluation control unit 39. For example, the answer target selection unit 41 can determine whether or not to instruct the text analysis unit 13 to execute an answer output process by comparing the WER value with a preset threshold value.
[0157] With the above configuration, the questionnaire processing system 1 according to the sixth embodiment obtains answers to questionnaire items from the text analysis unit 13 only when there is a high possibility that the speech recognition results are appropriate (i.e., the speech recognition results obtained from multiple speech recognition units match), thereby enabling stable acquisition of appropriate answers to questionnaire items.
[0158] Seventh embodiment Next, a questionnaire processing system 1 according to a seventh embodiment will be described with reference to Fig. 35. In the drawings and description of the seventh embodiment, components similar to those in any of the first to sixth embodiments described above are assigned the same reference numerals as those used in any of the first to sixth embodiments. Furthermore, the questionnaire processing system 1 according to the seventh embodiment is the same as that in any of the first to sixth embodiments, except for matters specifically mentioned below.
[0159] The questionnaire processing system 1 according to the seventh embodiment further includes a call terminal control unit 77. The main control unit 19 further includes an ACW (After Call Work) control unit 79.
[0160] The call terminal control unit 77 controls, via the call unit 5, which operator to assign to a customer inquiry (i.e., which operator terminal 2 to call) when a customer inquiry (i.e., an incoming call from the telephone 20) is received. At this time, an operator who is already on the phone with another customer should not be called. Furthermore, after the call ends, additional tasks, such as taking notes about the conversation, may be required. Dealing with customers can be mentally exhausting, and operators need to calm down and prepare for the next call. Therefore, an operator should not be called immediately after the call ends. Therefore, it is common to use an ACW time (a time when calls are not accepted after a call ends) to prevent operators from being called during or immediately after a call. However, the time required after a call ends actually varies depending on the content of the call. For example, if the customer who answered the call is angry, the operator will be under great stress, and it will take longer than usual for the operator to calm down and prepare for the next call. Therefore, it is desirable to change the ACW time depending on the content of the call.
[0161] The ACW control unit 79 can instruct the call terminal control unit 77 on the time until the operator starts a call with the next customer based on information regarding the call time and call duration (length of the call) obtained from the answer control unit 43 and the ACW time (the time during which work is stopped after one call is ended) set for each operator.
[0162] For example, when the answer control unit's response item for anger is "yes," the ACW control unit 79 can change the ACW time of the agent who handled the call to be longer than the reference value (i.e., to suspend further calls for a predetermined time), as shown in Fig. 24. This allows the operator who handled the call for which the alert was output to have a break, thereby reducing the psychological anxiety of the operator.
[0163] In the first to seventh embodiments, the questionnaire processing system 1 is applied to a questionnaire about operators who handle telephone inquiries from customers at a call center. However, the present invention is not limited to this, and the questionnaire processing system 1 may also be applied to a questionnaire about a teacher in an online lesson between one teacher and one student, as will be described below.
[0164] (Eighth embodiment) Next, a questionnaire processing system 1 according to an eighth embodiment will be described with reference to Fig. 36 and other figures. In the drawings and descriptions of the eighth embodiment, components similar to those of any of the first to seventh embodiments described above are assigned the same reference numerals as those used in any of the first to seventh embodiments. Furthermore, the questionnaire processing system 1 according to the eighth embodiment is the same as that of any of the first to seventh embodiments, except for matters specifically mentioned below.
[0165] The questionnaire processing system 1 according to the eighth embodiment has the same configuration as the questionnaire processing system 1 shown in Fig. 1. However, in the eighth embodiment, a teacher terminal used by a teacher and a student terminal used by a student are used instead of the operator terminal 2 and telephone set 20 shown in Fig. 1. In other words, in the eighth embodiment, the target of processing is not a call related to a customer's telephone inquiry at a call center as described above, but rather, for example, as shown in the speech recognition result in Fig. 36, the target of processing is not a call related to a customer's telephone inquiry at a call center as described above, but rather a voice (a conversation between a teacher and a student) in an online lesson between one teacher and one student.
[0166] The prompt conversion unit 33 converts the questionnaire items (before conversion) set by the person in charge of the questionnaire into question prompts (after conversion), as shown in FIG. 37, for example.
[0167] The question prompts include information such as the respondent's position (here, a fictitious elementary school boy taking online lessons) and the type of text being questioned (here, a transcript of a lesson).The question prompts also include questions about the teacher's response, the teacher's language, the teacher's understanding, the resolution of problems and questions that the student had, re-use (i.e., whether the student will use the online lesson again), the reasons for the student's dissatisfaction with the teacher's response, etc., and the student's opinions and impressions.
[0168] Furthermore, as a result of the answer output process, the text analysis unit 13 outputs answers (questionnaire answers) such as those shown in Fig. 38 to the answer control unit 43. The answers include answers to each question item in the question prompt.
[0169] Furthermore, the questionnaire processing system 1 may be applied to a questionnaire about a doctor in an online medical consultation between one doctor and one patient, for example, as will be described below.
[0170] (Ninth embodiment) Next, a questionnaire processing system 1 according to a ninth embodiment will be described with reference to Fig. 39 and other figures. In the drawings and descriptions relating to the ninth embodiment, components similar to those in any of the first to eighth embodiments described above are assigned the same reference numerals as those used in any of the first to eighth embodiments. Furthermore, the questionnaire processing system 1 according to the ninth embodiment is the same as that in any of the first to eighth embodiments, except for matters specifically mentioned below.
[0171] The questionnaire processing system 1 according to the eighth embodiment has the same configuration as the questionnaire processing system 1 shown in Fig. 1. However, in the ninth embodiment, a doctor terminal used by a doctor and a patient terminal used by a patient are used instead of the operator terminal 2 and telephone set 20 shown in Fig. 1. That is, in the ninth embodiment, the target of processing is not a call regarding a telephone inquiry from a customer at a call center as described above, but rather, for example, as shown in the speech recognition result in Fig. 39, the target of processing is not a call regarding a telephone inquiry from a customer at a call center as described above, but rather a voice (a conversation between a doctor and a patient) in an online medical consultation between one doctor and one patient.
[0172] The prompt conversion unit 33 converts the questionnaire items (before conversion) set by the person in charge of the questionnaire into question prompts (after conversion), as shown in FIG. 40, for example.
[0173] The question prompts include information such as the respondent's position (here, a hypothetical patient receiving online medical care) and the type of text being asked about (here, a transcript of the medical care).The question prompts also include questions about, for example, the doctor's response, the doctor's language, the doctor's understanding, the resolution of the patient's problems and doubts, re-use (i.e., whether the patient will use online medical care again), the patient's reasons for dissatisfaction with the doctor's response, etc., and the patient's opinions and impressions.
[0174] Furthermore, as a result of the answer output process, the text analysis unit 13 outputs answers (questionnaire answers) such as those shown in Fig. 41 to the answer control unit 43. The answers include answers to each question item in the question prompt.
[0175] As described above, the embodiments have been described as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these, and can be applied to embodiments in which modifications, substitutions, additions, omissions, etc. are made. Furthermore, it is also possible to combine the components described in the above embodiments to create new embodiments.
[0176] For example, in the above embodiment, an example was shown in which the survey processing system 1 is applied to calls made by telephone or equivalent communication means, but the target of the survey processing system 1 is not limited to calls. For example, the survey processing system 1 may omit voice recognition (processing that converts voice into text) and perform survey processing on conversations that have already been converted into text (for example, chat text). [Industrial Applicability]
[0177] The survey processing system, survey processing method, and survey processing program disclosed herein have the effect of being able to obtain the results of a survey regarding the response of an operator who spoke with a customer using a simple configuration, and are useful as a survey processing system, survey processing method, and survey processing program for obtaining the results of a survey. [Explanation of symbols]
[0178] 1: Survey processing system 2: Operator terminal 3: Operator terminal 5:Talking part 6: Comparative information accumulation 7: Audio receiving section 9: Voice recognition unit 11: Questionnaire setting section 13: Text analysis section 15: Analysis display section 17: Alert receiving section 19: Main control unit 20: Telephone 21: Display device 22: Input device 23: Microphone 24: Speaker 31: Questionnaire input section 33: Prompt conversion section 35: Audio input section 37: Voice recognition control unit 39: Voice recognition evaluation control unit 41: Respondent Selection Department 43: Answer control section 45: Answer storage section 47: Response collection section 49: Aggregated response storage unit 51: Analysis instruction input section 53: Data Analysis Department 55: Alert judgment unit 57: Alert summary control section 59: Alert section 60: Play button 63: Audio storage unit 65: Comparative information accumulation 67: Comparison result display area 69: Survey Decision Section 71: Text correction section 74: Hold music removal section 75: Recognition result synthesis section 77: Call terminal control unit 79: ACW control unit
Claims
1. A survey processing system for obtaining responses to a survey from the perspective of a customer who is a participant in a target conversation between a customer and an operator, comprising: one or more processors that execute a process for obtaining responses to the questionnaire; The one or more processors: Obtaining questionnaire items as questions regarding the questionnaire; Acquire the conversational voice of the subject; a voice recognition unit capable of executing voice recognition is instructed to recognize the conversational voice, and the conversational voice converted into text is obtained as a voice recognition result from the voice recognition unit; A questionnaire processing system that obtains answers to the questionnaire items from an answer output unit capable of outputting answers to questions in response to a text-converted conversational voice by instructing the answer output unit to output answers to the questionnaire items in response to the speech recognition results.
2. The one or more processors: a speech recognition evaluation unit capable of evaluating a result of speech recognition is instructed to evaluate the result of speech recognition, thereby obtaining the evaluation result from the speech recognition evaluation unit; The questionnaire processing system according to claim 1 , further comprising: determining whether or not to instruct the answer output unit to output the answer based on the evaluation result.
3. The questionnaire items are set based on a human input operation, the answer output unit is configured by a large-scale language model, The one or more processors:
2. The questionnaire processing system according to claim 1, wherein the questionnaire items are converted into prompts, and the prompts instruct the answer output unit to output answers to the questionnaire items in response to the speech recognition results.
4. the answer acquired from the answer output unit includes the customer's opinion and impression, The one or more processors: Sequentially accumulating the answers acquired from the answer output unit; The questionnaire processing system of claim 1, wherein the answer output unit is instructed to unify the spelling variations in the sentences of the accumulated answers, thereby obtaining from the answer output unit a plurality of answers including sentences with unified spelling variations.
5. The one or more processors: instructing a plurality of speech recognition units capable of evaluating speech recognition results to perform speech recognition on the conversational speech, and acquiring speech recognition results from the plurality of speech recognition units; determining whether the plurality of speech recognition results match; 2. The questionnaire processing system according to claim 1, wherein the answer output unit is instructed to output the answer only when it is determined that the plurality of speech recognition results match.
6. The one or more processors: determining whether an alert is necessary based on the answers to the questionnaire items acquired from the answer output unit; The questionnaire processing system according to claim 1 , wherein, when it is determined that the alert is necessary, an alert regarding the response to the questionnaire item is output to a preset alert receiving unit.
7. the conversation is an online conversation between a first user and a second user; the first user conducting a conversation with the second user via a first user terminal; The one or more processors: The questionnaire processing system according to claim 6, wherein when the alert is output, a terminal control unit that controls calls to the first user terminal is instructed to suspend further calls for a predetermined period of time.
8. The one or more processors: The questionnaire processing system according to claim 1 , wherein when the answer output unit is instructed to output an answer to the questionnaire item in the speech recognition result, the answer output unit is instructed to output a reason for the answer.
9. The one or more processors: if the conversation is an online conversation and includes music on hold, removing the music on hold; 2. The questionnaire processing system according to claim 1, wherein the voice recognition unit is instructed to perform voice recognition on the conversation voice from which the data related to the hold tone has been removed.
10. The one or more processors: Acquire a first voice related to a voice of a first user and a second voice related to a voice of a second user; acquiring the first speech converted into text and the second speech converted into text from the speech recognition unit as speech recognition results, respectively; The survey processing system of claim 1, wherein the speech recognition result for instructing the answer output unit to output the answer to the survey item includes a distinguishable voice of the first user and a distinguishable voice of the second user.
11. A survey processing method for obtaining responses to a survey from the perspective of a customer who is a participant in a target conversation between a customer and an operator, comprising: One or more computers Obtaining questionnaire items as questions regarding the questionnaire; Acquire the conversational voice of the subject; a voice recognition unit capable of executing voice recognition is instructed to recognize the conversational voice, and the conversational voice converted into text is obtained as a voice recognition result from the voice recognition unit; A questionnaire processing method, comprising: instructing an answer output unit capable of outputting answers to questions in response to a conversational voice that has been converted into text to output answers to the questionnaire items in response to the speech recognition results, thereby obtaining answers to the questionnaire items from the answer output unit.
12. A questionnaire processing program that causes a computer to execute information processing for obtaining responses to a questionnaire from the perspective of a customer who is a participant in a target conversation between a customer and an operator, comprising: The information processing includes: Obtaining questionnaire items as questions regarding the questionnaire; Acquire the conversational voice of the subject; a voice recognition unit capable of executing voice recognition is instructed to recognize the conversational voice, and the conversational voice converted into text is obtained as a voice recognition result from the voice recognition unit; A questionnaire processing program that includes a procedure for acquiring answers to the questionnaire items from an answer output unit that is capable of outputting answers to questions in response to a text-converted conversational voice, by instructing the answer output unit to output the answers to the questionnaire items in the speech recognition results.
Citation Information
Patent Citations
Call content managing device and program thereof
JP2004096149A
Investigation system using voice response
JP2004229014A
Call center followup processing system and followup processing method
JP2017204868A
Matching support system
JP2021182290A
System, program, and method for surveying
JP2022125096A