Survey processing system, survey processing method, and survey processing program

The questionnaire processing system effectively captures operator responses through speech recognition and analysis, addressing data limitations and burden issues by providing accurate and efficient survey results with reduced operator stress.

JP2026067994APending Publication Date: 2026-04-21PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2026-02-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Conventional survey systems fail to accurately capture operator responses during customer inquiries due to limited training data and impractical direct customer surveys, placing a heavy burden on both operators and customers.

Method used

A questionnaire processing system that utilizes speech recognition to transcribe and analyze call audio, providing answers to survey questions based on high-accuracy speech recognition results, and outputs responses through a large-scale language model while managing inconsistencies and alerting necessary actions.

Benefits of technology

Enables accurate and efficient acquisition of operator responses during phone calls with reduced psychological anxiety for operators by ensuring high-quality speech recognition and standardized feedback aggregation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026067994000001_ABST
    Figure 2026067994000001_ABST
Patent Text Reader

Abstract

The results of surveys regarding services provided by service providers via two-way voice calls with customers, such as the responses of operators who answered customer calls, are obtained using a simple configuration. [Solution] The questionnaire processing system 1 comprises one or more processors that perform processing to obtain the results of the questionnaire. The one or more processors obtain questionnaire items as questions related to the questionnaire, obtain call audio between the customer and the operator, and instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the call audio data. The system then obtains the call audio data converted into text from the speech recognition unit as a result of speech recognition. The system then instructs an answer output unit capable of outputting answers to questions on the converted call audio to output answers to the questionnaire items as questions on the speech recognition results. The system is configured to obtain the answers to the questionnaire items from the answer output unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , , , , ,

[0006] , , , ,

[0005] , , , , , , ,

[0003] ,

[0001] The present disclosure relates to a questionnaire processing system, a questionnaire processing method, and a questionnaire processing program for obtaining questionnaire results.

Background Art

[0002] Conventionally, an administrator such as a call center may conduct a questionnaire for a customer who made a call to determine whether the operator's response to the customer's inquiry call was appropriate. Generally, in conducting a questionnaire, responses are obtained from a plurality of customers, and the responses are aggregated.

[0003] Also, a system that uses a plurality of personalized AIs (Artificial Intelligence) to conduct a questionnaire is known in order to reduce the workload of questionnaire surveys (see Patent Document 1). According to this conventional system, for example, it is possible to obtain a response from a personalized AI for a question such as "Do you support Mr. XX in the next presidential election?".

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

[0007] Therefore, the main purpose of this disclosure is to provide a survey processing system, a survey processing method, and a survey processing program that can obtain the results of surveys regarding services provided via two-way voice calls between the service provider and the customer, including the responses of operators who have spoken with customers, using a simple configuration. [Means for solving the problem]

[0008] The questionnaire processing system of this disclosure is a questionnaire processing system for obtaining the results of a questionnaire regarding an operator who has spoken with a customer, and comprises one or more processors that perform processing for obtaining the results of the questionnaire, the one or more processors obtain questionnaire items as questions related to the questionnaire, obtain the audio of a conversation between the customer and the operator, instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the audio, thereby obtaining the audio from the speech recognition unit as text as a result of speech recognition, and instruct an answer output unit capable of outputting answers to questions on the text of the audio to output answers to the questionnaire items on the speech recognition results, thereby obtaining answers to the questionnaire items from the answer output unit.

[0009] Furthermore, the questionnaire processing method of this disclosure is a questionnaire processing method for obtaining the results of a questionnaire regarding an operator who has spoken with a customer, wherein one or more computers acquire questionnaire items as questions related to the questionnaire, acquire call audio data relating to a call between the customer and the operator, instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the call audio, thereby obtaining the call audio transcribed into text as a result of speech recognition from the speech recognition unit, and instruct an answer output unit capable of outputting answers to questions related to the transcribed call to output answers to the questionnaire items based on the speech recognition results, thereby obtaining answers to the questionnaire items from the answer output unit.

[0010] Furthermore, the questionnaire processing program of this disclosure is a questionnaire processing program that causes a computer to perform information processing to obtain the results of a questionnaire regarding an operator who has spoken with a customer, and comprises one or more processors that perform processing to obtain the results of the questionnaire, wherein the information processing includes a procedure to obtain questionnaire items as questions related to the questionnaire, obtain call audio data relating to a conversation between the customer and the operator, instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the call audio to obtain the call audio transcribed into text as a result of speech recognition from the speech recognition unit, and instruct an answer output unit capable of outputting answers to the questions in the transcribed call to output answers to the questionnaire items in the speech recognition results to obtain answers to the questionnaire items from the answer output unit. [Effects of the Invention]

[0011] According to this disclosure, the results of a survey regarding services provided by the service provider via two-way voice calls with the customer, such as the response of the operator who spoke with the customer, can be obtained with a simple configuration. [Brief explanation of the drawing]

[0012] [Figure 1] Overall configuration diagram of the questionnaire processing system according to the first embodiment [Figure 2] Block diagram showing the schematic configuration of the operator terminal shown in Figure 1. [Figure 3] Flowchart showing the flow of questionnaire processing performed in the questionnaire processing system according to the first embodiment. [Figure 4] Figure 3 shows an example of the conversion from questionnaire items to question prompts in step ST101. [Figure 5] An explanatory diagram showing an example of a voice evaluation prompt in step ST104 in Figure 3. [Figure 6] Figure 3 shows an explanatory diagram illustrating an example of the evaluation results regarding the speech recognition results in step ST104. [Figure 7] Figure 3 is an explanatory diagram showing an example of the speech recognition result input from the response control unit to the text analysis unit in step ST106. [Figure 8] Figure 3 shows an example of a response (questionnaire response) output from the text analysis unit to the response control unit in step ST106. [Figure 9] Figure 3 shows an example of the response information stored in the response storage unit during step ST108. [Figure 10] This explanation shows an example of a response aggregation prompt that is input from the response aggregation unit to the text analysis unit in step ST109 in Figure 3. [Figure 11] Figure 3 shows an explanatory diagram illustrating an example of specific opinions and feedback entered from the response aggregation unit to the text analysis unit in step ST109 of Figure 3. [Figure 12] This diagram in Figure 3 illustrates an example of the result of unifying variations in notation, which is output from the text analysis unit to the response aggregation unit in step ST109. [Figure 13] Figure 3 shows an example of aggregated responses stored in the aggregated response storage unit in step ST109. [Figure 14] Figure 3 shows an example of the data analysis results in step ST110. [Figure 15] Explanatory diagram showing an example of a response determined by the alert determination unit 55 to require an alert in step ST111 in FIG. 3 [Figure 16] Explanatory diagram showing an example of a voice recognition result determined by the alert determination unit 55 to require an alert in step ST111 in FIG. 3 [Figure 17] Explanatory diagram showing an example of a summary prompt input from the alert summary control unit to the text analysis unit in step ST111 in FIG. 3 [Figure 18] Explanatory diagram showing an example of a response (summary result) output from the text analysis unit to the alert summary control unit in step ST111 in FIG. 3 [Figure 19] Explanatory diagram showing an example of an alert notification output from the alert unit to the alert reception unit in step ST111 in FIG. 3 [Figure 20] Explanatory diagram showing another example of a voice recognition result input from the response control unit to the text analysis unit in step ST106 in FIG. 3 [Figure 21] Explanation showing another example of a question prompt input from the response aggregation unit to the text analysis unit in step ST106 in FIG. 3 [Figure 22] Explanatory diagram showing an example of a determination rule for a notification destination based on an alert condition in step ST107 in FIG. 3 [Figure 23] Explanatory diagram showing an example of an alert notification output from the alert unit to the alert reception unit in step ST111 in FIG. 3 [Figure 24] Explanatory diagram showing an example of a response (explanation of problem location) output from the text analysis unit to the response control unit [Figure 25] Overall configuration diagram of the questionnaire processing system according to the second embodiment [Figure 26] Flow diagram showing the flow of questionnaire change processing performed by the questionnaire processing system according to the second embodiment [Figure 27] Explanatory diagram showing an example of a comparison response set [Figure 28] Overall configuration diagram of the questionnaire processing system according to the third embodiment [Figure 29] Figure 28 is an explanatory diagram showing an example of a correction prompt input from the text correction unit to the text analysis unit. [Figure 30] Figure 28 is an explanatory diagram showing an example of the speech recognition results (before correction) input from the text correction unit to the text analysis unit and the speech recognition results (after correction) output from the text analysis unit to the recognition result correction unit. [Figure 31] Overall configuration diagram of the questionnaire processing system according to the fourth embodiment [Figure 32] Overall configuration diagram of the questionnaire processing system according to the fifth embodiment [Figure 33] Figure 32 is an explanatory diagram showing an example of speech recognition results synthesized by the recognition result synthesis unit shown in Figure 32. [Figure 34] Overall configuration diagram of the questionnaire processing system according to the sixth embodiment [Figure 35] Overall configuration diagram of the questionnaire processing system according to the 7th embodiment [Figure 36] An explanatory diagram showing an example of the speech recognition result by the speech recognition unit according to the eighth embodiment. [Figure 37] This diagram illustrates an example of the conversion of questionnaire items to question prompts by the prompt conversion unit according to the eighth embodiment. [Figure 38] This diagram illustrates an example of a response output from the text analysis unit to the response control unit according to the eighth embodiment. [Figure 39] An explanatory diagram showing an example of the speech recognition result by the speech recognition unit according to the ninth embodiment. [Figure 40] This diagram illustrates an example of the conversion of questionnaire items to question prompts by the prompt conversion unit according to the ninth embodiment. [Figure 41] This diagram illustrates an example of a response output from the text analysis unit to the response control unit according to the 9th embodiment. [Modes for carrying out the invention]

[0013] The first invention made to solve the aforementioned problems is an questionnaire processing system for obtaining the results of a questionnaire regarding an operator who has spoken with a customer, comprising one or more processors that perform processing for obtaining the results of the questionnaire, wherein the one or more processors obtain questionnaire items as questions related to the questionnaire, obtain the audio of the conversation between the customer and the operator, instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the audio, thereby obtaining the audio from the speech recognition unit as text as a result of speech recognition, and instruct an answer output unit capable of outputting answers to the questions on the text of the audio to output answers to the questionnaire items on the speech recognition results, thereby obtaining answers to the questionnaire items from the answer output unit.

[0014] According to this, the results of a survey regarding the customer service provided by operators during phone calls can be obtained using a simple configuration.

[0015] Furthermore, the second invention is configured such that the one or more processors instruct a speech recognition evaluation unit capable of evaluating the speech recognition result to evaluate the speech recognition result, obtain the evaluation result from the speech recognition evaluation unit, and determine whether or not to instruct the response output unit to output the response based on the evaluation result.

[0016] According to this, it becomes possible to exclude low-rated speech recognition results based on the evaluation results (i.e., not instruct the output of a response for low-rated speech recognition results), so that appropriate answers to the questionnaire items can be obtained from the response output unit based only on call audio with high speech recognition accuracy.

[0017] Furthermore, the third invention is configured such that the questionnaire items are set based on human input operations, the response output unit is composed of a large-scale language model, and the one or more processors convert the questionnaire items into prompts, and the prompts instruct the response output unit to output answers to the questionnaire items in the speech recognition results.

[0018] According to this, based on the prompts into which the survey items have been converted, the appropriate answers to the survey items can be obtained from the response output unit.

[0019] Furthermore, the fourth invention is configured such that the answers obtained from the answer output unit include the customer's opinions and impressions, and the one or more processors sequentially store the answers obtained from the answer output unit and instruct the answer output unit to unify inconsistencies in notation in the stored multiple answers, thereby obtaining multiple answers from the answer output unit that include sentences with unified notation.

[0020] According to this, users of the survey responses (for example, those who manage operators) can easily grasp customers' opinions and feedback on their calls with operators, based on customer opinions and feedback where inconsistencies in wording in the responses are standardized.

[0021] Furthermore, the fifth invention is configured such that the one or more processors instruct a plurality of speech recognition units capable of evaluating the results of speech recognition to perform speech recognition on the call audio, thereby obtaining the speech recognition results from each of the plurality of speech recognition units, determining whether the plurality of speech recognition results match, and only if it is determined that the plurality of speech recognition results match, instruct the response output unit to output the response.

[0022] According to this, responses to survey items are retrieved from the response output unit only when there is a high probability that the speech recognition results are appropriate (i.e., when the speech recognition results obtained from multiple speech recognition units match), thus enabling the stable acquisition of appropriate responses to survey items.

[0023] Furthermore, the sixth invention is configured such that the one or more processors determine whether an alert is necessary based on the answers to the questionnaire items obtained from the answer output unit, and if it is determined that an alert is necessary, it outputs an alert regarding the answers to the questionnaire items to a pre-configured alert receiving unit.

[0024] According to this, users of the survey responses (for example, the administrator of operators using the alert receiving unit) can easily identify responses that require confirmation (for example, responses regarding calls with operators that customers were dissatisfied with) by receiving alerts.

[0025] Furthermore, the seventh invention is configured such that the operator makes calls with the customer using an operator terminal, and when the alert is output, the one or more processors instruct the call terminal control unit, which controls calls to the operator terminal, to suspend further calls for a predetermined time.

[0026] According to this, by giving the operator who made the call that triggered the alert a rest (i.e., suspending calls to the operator's terminal for a predetermined period of time), the operator's psychological anxiety can be reduced.

[0027] Furthermore, the eighth invention is configured such that when the one or more processors instruct the response output unit to output the response to the questionnaire item in the speech recognition result, they also instruct the processor to output the basis for the response.

[0028] According to this, users of survey responses (for example, those managing operators) can easily understand the customer's intent behind the responses based on the reasoning behind them.

[0029] Furthermore, the ninth invention is configured such that, if the call between the customer and the operator includes telephone hold music, the one or more processors remove the hold music from the call audio and instruct the speech recognition unit to perform speech recognition on the call audio from which the data related to the hold music has been removed.

[0030] According to this, by eliminating the hold music included in the call audio, it is possible to avoid the hold music negatively affecting the response output from the response output unit.

[0031] Furthermore, the tenth invention is configured such that the one or more processors acquire customer voice and operator voice as the call audio, respectively, and acquire the customer voice and the operator voice transcribed into text from the speech recognition unit as the speech recognition results, respectively, and the speech recognition results for instructing the response output unit to output the responses to the questionnaire items include the customer voice and the operator voice in an identifiable manner.

[0032] According to this, appropriate answers to questionnaire items can be obtained from the response output unit based on call audio data in which speakers have been distinguished (i.e., speech recognition results).

[0033] Furthermore, the 11th invention is a questionnaire processing method for obtaining the results of a questionnaire regarding an operator who has spoken with a customer, wherein one or more computers obtain questionnaire items as questions related to the questionnaire, obtain call audio data relating to the conversation between the customer and the operator, instruct a speech recognition unit capable of performing speech recognition to perform speech recognition on the call audio, thereby obtaining the call audio transcribed into text as a result of speech recognition from the speech recognition unit, and instruct an answer output unit capable of outputting answers to the questions related to the transcribed call to output answers to the questionnaire items based on the speech recognition results, thereby obtaining answers to the questionnaire items from the answer output unit.

[0034] According to this, the results of a survey regarding the customer service provided by operators during phone calls can be obtained using a simple configuration.

[0035] Furthermore, the twelfth invention is a survey processing program that causes a computer to perform information processing to obtain the results of a survey concerning an operator who has spoken with a customer, comprising one or more processors that perform processing to obtain the results of the survey, wherein the information processing includes the steps of obtaining survey items as questions related to the survey, obtaining call audio data relating to a conversation between the customer and the operator, instructing a speech recognition unit capable of performing speech recognition to perform speech recognition on the call audio to obtain the call audio transcribed into text as a result of speech recognition from the speech recognition unit, and instructing an answer output unit capable of outputting answers to the questions in the transcribed call to output answers to the survey items in relation to the speech recognition results to obtain answers to the survey items from the answer output unit.

[0036] According to this, the results of a survey regarding the customer service provided by operators during phone calls can be obtained using a simple configuration.

[0037] The embodiments of this disclosure will be described below with reference to the drawings.

[0038] (First Embodiment) As shown in Figure 1, the questionnaire processing system 1 according to the first embodiment includes an operator terminal 3, a call unit 5, a voice receiving unit 7, a voice recognition unit 9, a questionnaire setting unit 11, a text analysis unit 13, an analysis display unit 15, an alert receiving unit 17, and a main control unit 19.

[0039] The survey processing system 1 is used, for example, to obtain the results of a survey regarding operators who have answered customer inquiries over the phone at a call center. The survey processing system 1 can obtain answers to the survey questions not from actual customers, but from a virtual customer that simulates the customer who made the call. Customers can speak with operators by calling the call center using telephone 20.

[0040] Note that Figure 1 shows only one operator terminal 3 used by one operator (not shown). However, typically a call center has multiple operators, and each operator is assigned an operator terminal 3. Also, Figure 1 shows only one telephone 20 used by one customer (not shown). However, typically inquiries to a call center are made through telephones used by multiple customers.

[0041] The operator terminal 3 consists of a PC, smartphone, tablet, or the like. However, the operator terminal 3 is not limited to these and can be composed of any device that has at least a call function.

[0042] As shown in Figure 2, the operator terminal 3 includes a display device 21 such as an LCD panel, input devices 22 such as a touch panel, keyboard, and mouse, a microphone 23, and a speaker 24. The microphone 23 and speaker 24 may be a headset. A device with the same configuration as the operator terminal 3 may be used as the telephone 20.

[0043] The operator can make calls to customers using the telephone 20 using the microphone 23 and speaker 24 of the operator terminal 3. The operator terminal 3 uses an application such as a browser to display various screens that the operator should view on the display device 21 based on the display information transmitted from the call unit 5. The operator terminal 3 can detect screen operations performed by the operator using the input device 22 and transmit that operation information to the call unit 5.

[0044] The call unit 5 relays the transmission and reception of voice data between the customer's telephone 20 and the operator's terminal 3 used by the operator, thereby facilitating communication between the customer and the operator. The call unit 5 can be implemented, for example, by a well-known cloud service that provides cloud contact center functionality. Alternatively, the call unit 5 may be implemented by a telephone system that utilizes business phone functionality via an internet connection, by building a "PBX (private branch exchange)" function on the cloud.

[0045] The voice receiving unit 7 receives (collects) the audio of the conversation between the customer and the operator that takes place in the call unit 5 (hereinafter sometimes simply referred to as "call audio"). The voice receiving unit 7 is implemented, for example, by a known cloud service that provides a call recording function. The voice receiving unit 7 may also be implemented as part of a known cloud contact center function.

[0046] The speech recognition unit 9 performs a process (hereinafter referred to as "speech recognition processing") to convert the call audio acquired by the voice receiving unit 7 into text using speech recognition, in response to instructions from the main control unit 19. The speech recognition unit 9 is implemented, for example, by a known cloud service that provides a function to convert speech to text.

[0047] The survey setting unit 11 consists of a PC, smartphone, or tablet device. The survey setting unit 11 is used by a survey manager, for example, belonging to a customer management department. The survey manager can set up survey questions about the operator who spoke with the customer by accessing the main control unit 19 from the survey setting unit 11. The set questions are sent to the main control unit 19.

[0048] The text analysis unit 13 is implemented using a well-known cloud service that utilizes a large language model (LLM), such as Google Gemini®, Microsoft Bing AI®, or ChatGPT. The text analysis unit 13 can perform various processes in response to instructions from the main control unit 19. The multiple processes performed by the text analysis unit 13 may each be implemented using multiple different cloud services. Alternatively, the large language model may be implemented on an on-premises server instead of using a cloud service.

[0049] For example, the text analysis unit 13 (an example of a speech recognition evaluation unit) can execute a process (hereinafter referred to as "speech recognition evaluation process") that outputs an evaluation of the speech recognition result of the speech recognition unit 9 (i.e., the call audio converted into text by the speech recognition process).

[0050] Furthermore, the text analysis unit 13 (an example of an answer output unit) can perform a process (hereinafter referred to as "answer output process") to output answers to pre-prepared questions based on the speech recognition results of the speech recognition unit 9. The pre-prepared questions include questions related to the questionnaire set by the questionnaire setting unit 11 (hereinafter referred to as "questionnaire items" as needed). The purpose of the questionnaire is to have customers (more precisely, virtual customers realized by the text analysis unit 13) answer questions such as their satisfaction level and dissatisfaction with their conversations with operators. In addition, the text analysis unit 13 can perform a process (hereinafter referred to as "answer aggregation process") to unify variations in text notation for multiple answers to questionnaire items.

[0051] Furthermore, as will be described later, the text analysis unit 13 can perform a process to generate a summary of the corresponding call audio (hereinafter referred to as the "alert summary generation process") when it becomes necessary to output an alert to the operator's manager (for example, a supervisor) regarding the responses obtained to the questionnaire items.

[0052] The analysis display unit 15 is comprised of a PC, smartphone, or tablet device. The analysis display unit 15 is used by an analyst responsible for analyzing survey results, such as an operator manager. In the analysis display unit 15, the analyst can give instructions to the main control unit 19 regarding data analysis of the accumulated responses. Such data analysis includes, for example, statistical processing of the accumulated responses and the creation of tables and graphs. The analyst can also instruct the main control unit 19 to display the accumulated responses and the results of their data analysis.

[0053] The alert receiving unit 17 consists of a PC, smartphone, or tablet device. The alert receiving unit 17 is used by a troubleshooter, such as an operator manager. The troubleshooter can receive alerts output from the main control unit 19 using the alert receiving unit 17. This allows the troubleshooter to address customer complaints regarding calls with operators and provide support to the operators based on the content of the alerts.

[0054] The main control unit 19 includes a questionnaire input unit 31, a prompt conversion unit 33, a voice input unit 35, a voice recognition control unit 37, a voice recognition evaluation control unit 39, a response target selection unit 41, a response control unit 43, a response storage unit 45, a response aggregation unit 47, an aggregated response storage unit 49, an analysis instruction input unit 51, a data analysis unit 53, an alert determination unit 55, an alert summary control unit 57, and an alert unit 59. The main control unit 19 is configured to communicate with the voice receiving unit 7, a voice recognition unit 9, a text analysis unit 13, a questionnaire setting unit 11, an analysis display unit 15, and an alert receiving unit 17, etc., via a known communication network.

[0055] The survey input unit 31 receives input related to the survey from the survey setting unit 11 (an example of human input operation). The survey items to be entered may be changed as appropriate by the person in charge of the survey.

[0056] Prompt input is required to give instructions to the large-scale language model. However, survey administrators generally do not know what to write as a prompt. Therefore, the prompt conversion unit 33 converts the survey items input from the survey setting unit 11 into prompts (hereinafter referred to as "question prompts" as needed). The question prompts include strings that clearly convey specific tasks and requirements to the text analysis unit 13. The converted prompts are also input to the answer control unit 43. The converted prompts may include a pair of provisional speech recognition results and example answers to realize few-shot learning.

[0057] The voice input unit 35 receives call audio (voice data) from the voice receiving unit 7. For example, when a call between a customer and an operator (i.e., one call between the telephone 20 and the operator terminal 3) ends, the voice input unit 35 receives the call audio from the voice receiving unit 7. In order to increase the processing speed of voice recognition, the call audio may be input sequentially from the voice receiving unit 7.

[0058] The speech recognition control unit 37 acquires the call audio input to the voice input unit 35 and instructs the speech recognition unit 9 to perform speech recognition processing on that call audio. As a result, the speech recognition control unit 37 acquires the call audio converted into text (i.e., text data of the call audio) as the result of the speech recognition processing by the speech recognition unit 9 (hereinafter sometimes simply referred to as "speech recognition result"). In order to increase the processing speed of speech recognition, the speech recognition unit 9 may be instructed to perform speech recognition processing on the call audio that is input sequentially, and the speech recognition results may be combined for each call.

[0059] Furthermore, the voice recognition control unit 37 can acquire information about noisy lines from the call unit 5, such as calls from mobile phone numbers, and exclude the audio of calls using such lines from voice recognition processing. In addition, the voice recognition control unit 37 can acquire information about wrong numbers from the call unit 5, such as calls from phone numbers not registered in advance, and exclude the audio of such wrong numbers from voice recognition processing.

[0060] The speech recognition evaluation control unit 39 instructs the text analysis unit 13 to execute speech recognition evaluation processing on the speech recognition results obtained by the speech recognition control unit 37. As a result, the speech recognition evaluation control unit 39 obtains the evaluation results (i.e., the results of the speech recognition evaluation processing) regarding the speech recognition results from the speech recognition unit 9 from the text analysis unit 13. These evaluation results serve as an indicator of the quality of the text of the call audio obtained by the speech recognition processing (in this case, whether or not it is suitable as a target for response output processing by the text analysis unit 13).

[0061] The response target selection unit 41 determines, based on the evaluation results regarding the speech recognition results obtained by the speech recognition evaluation control unit 39, whether or not to instruct the text analysis unit 13 to execute the response output processing (i.e., output the response) for the speech recognition results to be evaluated (i.e., the text of the call audio). If the content of the call audio is not accurately reflected in the speech recognition results (i.e., the evaluation of the speech recognition results is low), it may negatively affect the results of the response output processing and, consequently, the data analysis results. Therefore, the response target selection unit 41 excludes speech recognition results with low evaluations from the target of the response output processing (i.e., it determines that it should not instruct the text analysis unit 13 to execute the response output processing).

[0062] The response control unit 43 acquires the speech recognition result that the response target selection unit 41 has determined should be instructed to execute the response output processing on the text analysis unit 13. The response control unit 43 also acquires a prompt corresponding to the speech recognition result from the prompt conversion unit 33. Using the acquired speech recognition result and the corresponding prompt, the response control unit 43 instructs the text analysis unit 13 to execute the response output processing (i.e., output the answers to the questionnaire items). As a result, the response control unit 43 acquires the answers to the questionnaire items regarding the speech recognition result from the text analysis unit 13. Note that the answers to the questionnaire items may be output in a specific language (e.g., Japanese) regardless of the language of the call between the customer and the operator or the language used for the prompts, by instructing the text analysis unit 13 to output in a specific language.

[0063] The response information acquired by the response control unit 43 is sequentially stored in the response storage unit 45. The response storage unit 45 includes, for example, a database built on the cloud.

[0064] The response aggregation unit 47 instructs the text analysis unit 13 to perform response aggregation processing on the response information stored in the response storage unit 45. As a result, the response aggregation unit 47 obtains multiple responses that have undergone response aggregation processing (hereinafter referred to as "aggregated responses") from the text analysis unit 13.

[0065] The aggregated responses include multiple responses with unified text formatting for free-response type answers such as customer opinions and feedback. For example, the response aggregation unit 47 can instruct the text analysis unit 13 to perform the response aggregation process when the main control unit 19 (in this case, the analysis instruction input unit 51) receives an instruction regarding data analysis from the analysis display unit 15. Alternatively, the response aggregation unit 47 may instruct the text analysis unit 13 to perform the response aggregation process when a predetermined amount of response information has been accumulated in the response storage unit 45. The response aggregation unit 47 may also instruct the text analysis unit 13 to perform the response aggregation process at a specific scheduled time, such as 0:00 on the 1st of each month.

[0066] The aggregated response information obtained by the response aggregation unit 47 is sequentially stored in the aggregated response storage unit 49. The aggregated response storage unit 49 includes, for example, a database built on the cloud.

[0067] The analysis instruction input unit 51 receives instructions regarding data analysis from the analysis display unit 15. The input instructions regarding data analysis may be modified as appropriate by the analyst.

[0068] The data analysis unit 53 performs data analysis on the aggregated responses stored in the aggregated response storage unit 49 based on data analysis instructions entered in the analysis instruction input unit 51. As part of the data analysis, the data analysis unit 53 can process the aggregated response data using statistical methods and display the results as tables or graphs using visualization methods. The data analysis unit 53 also displays the results of such data analysis on the analysis display unit 15. For example, the data analysis unit 53 can use a cloud-based BI (Business Intelligence) tool to perform data analysis and display the results on the analysis display unit 15.

[0069] The alert determination unit 55 determines whether an alert is necessary based on the responses to the survey items obtained by the response control unit 43. For example, the alert determination unit 55 can determine that an alert is necessary if the responses to the survey items include information indicating customer dissatisfaction.

[0070] The alert summary control unit 57 instructs the text analysis unit 13 to execute the alert summary generation process for responses that the alert determination unit 55 has determined require an alert. This allows the alert summary control unit 57 to obtain summary information (hereinafter referred to as "alert information" as needed) related to the corresponding call audio from the text analysis unit 13. The alert information includes, for example, the name of the store used by the customer (or the name of the product purchased or service received by the customer), customer identification information (e.g., the customer's name), operator identification information (e.g., the operator's name), the content of the customer's inquiry, and the content of the operator's response. If the call unit 5 is configured with a cloud service PBX, customer identification information and operator identification information may be managed by the cloud PBX. In that case, customer and operator identification information may be obtained directly from the voice receiving unit 7.

[0071] If the alert determination unit 55 determines that an alert is necessary, the alert unit 59 outputs (or transmits) the alert information acquired by the alert summary control unit 57 to the alert receiving unit 17. This allows the troubleshooting personnel using the alert receiving unit 17 to check the content of the alert. However, the alert unit 59 may output the response indicating that an alert is necessary to the alert receiving unit 17 without adding a summary. In that case, the alert summary control unit 57 may be omitted in the main control unit 19.

[0072] At least some of the functions of each part of the main control unit 19 described above can be realized, for example, by cloud computing (i.e., servers, storage, network infrastructure, databases, and software in a cloud environment). The main functions of the main control unit 19 can be realized by one or more processors in one or more servers executing a predetermined control program. In addition, at least some of the functions of each part of the main control unit 19 may be realized by edge computing (edge ​​servers, etc.).

[0073] Next, with reference to Figure 3 and other figures, the flow of the questionnaire processing performed in the questionnaire processing system 1 according to the first embodiment will be described. Here, the questionnaire processing is the process of obtaining the results of a questionnaire regarding the operator who spoke with the customer (including the customer's responses and the results of data analysis of those responses).

[0074] In the survey processing system 1, when survey items are first entered from the survey setting unit 11 to the survey input unit 31, the prompt conversion unit 33 converts those survey items into question prompts (ST101).

[0075] In step ST101, for example, as shown in Figure 4, the survey items (before conversion) set by the survey administrator are converted into question prompts (after conversion).

[0076] Question prompts include information such as the respondent's role (in this case, a hypothetical customer) and the type of text being questioned (in this case, a call center transcript). They also include questions about the operator's response, language, understanding, resolution of customer problems or questions, likelihood of repeat use (i.e., whether the customer is likely to become a repeat customer), reasons for customer dissatisfaction with the operator's response (an example of the basis for the answer), and customer opinions and feedback.

[0077] For example, as shown in Figure 4, the prompt begins with the following text: "You are a customer who has called the call center. You will now enter a transcript of the call center conversation. Please answer the survey questions from the customer's perspective." This allows the user to virtually answer the survey from the perspective of the customer who made the call.

[0078] Next, add the following prompt: "Please output in JSON format using the following key." This will allow the response to be output in a JSON format that is easy for the computer to handle.

[0079] Next, add "# Key, Question" to the prompt. This instructs you to write the key text in JSON format and the question text in a corresponding format.

[0080] Furthermore, key phrases and question content are extracted from the original survey items and included in the prompt. For example, Q1. Response How was your experience with the call center? Choose from "Satisfied", "Mostly Satisfied", "Average", "Somewhat Dissatisfied", or "Dissatisfied". Please answer. For the survey item, the string after QX. is extracted as the Key, and the string after the newline following the Key is recognized as the question, converted as shown below, and appended to the prompt. How was your experience with the customer service and the call center? Please choose your answer from the following options: "Satisfied," "Mostly satisfied," "Neutral," "Somewhat dissatisfied," or "Dissatisfied."

[0081] Furthermore, the question prompts may include, for example, actual customer attribute information (age, gender, etc.) contained in the call audio. Also, if the call unit 5 is configured with a cloud service PBX, customer identification information may be managed by the cloud PBX, and the prompt conversion unit 33 may be able to obtain the identification information. For example, if the customer who made the call was a "male under 10 years old," the prompt may be converted to match the customer's identification information, such as "You are a male customer under 10 years old who called the call center. You will now enter the transcript of the call center. Please answer the survey questions from the customer's perspective." In this way, by instructing the text analysis unit to have a role that includes the attribute information of the customer who made the call, it is possible to obtain survey responses that more closely resemble the customer who made the call. In addition, the question prompts should be adjusted so that even if the same questions are asked to the actual customer who made the call audio, the same answers as the virtual customer (i.e., the text analysis unit 13) can be obtained. Furthermore, the question prompts may be set to require specific scores to be answered for each question.

[0082] Note that once step ST101 has been executed, it can be omitted until the survey items are changed (i.e., the same question prompt is used repeatedly).

[0083] Once a call between the customer and the operator has ended (Yes in ST102), the call audio is input from the audio receiving unit 7 to the audio input unit 35, and the speech recognition control unit 37 obtains the speech recognition result of the call audio from the speech recognition unit 9 (ST103).

[0084] Next, the speech recognition evaluation control unit 39 obtains the evaluation results regarding the speech recognition results acquired in step ST103 from the text analysis unit 13 (ST104).

[0085] In step ST104, a prompt such as the one shown in Figure 5 (hereinafter referred to as the "voice evaluation prompt") is input from the voice recognition evaluation control unit 39 to the text analysis unit 13 along with the voice recognition result.

[0086] The prompts for voice evaluation include information such as the respondent's role (in this case, a hypothetical professional editor) and the type of text being questioned (in this case, a transcript of a conversation generated by speech recognition). The prompts for questions include questions related to the accuracy of the speech recognition, the reasons for the professional editor's dissatisfaction with that accuracy, and the professional editor's opinions and impressions.

[0087] Furthermore, in step ST104, as shown in Figure 6 for example, the evaluation results regarding the speech recognition results are output from the text analysis unit 13 to the speech recognition evaluation control unit 39. These evaluation results include the answers to each question in the speech evaluation prompt.

[0088] Next, the response selection unit 41 determines whether the speech recognition result is appropriate as a response based on the evaluation results obtained by the speech recognition evaluation control unit 39, and excludes inappropriate speech recognition results from the response output processing (ST105). For example, as shown in Figure 6, speech recognition results in which the "Speech Recognition Accuracy" item of the response is "Somewhat dissatisfied" or "Dissatisfied" are judged as inappropriate and excluded. As a result, only highly rated speech recognition results are selected as targets for response output processing.

[0089] Subsequently, the response control unit 43 uses the speech recognition results and corresponding question prompts determined to be appropriate by the response target selection unit 41 to obtain responses to the questionnaire items regarding the speech recognition results from the text analysis unit 13 (ST106).

[0090] In step ST106, the appropriate speech recognition result (i.e., the text to be answered) and the question prompt, as shown in Figure 7, are input from the answer control unit 43 to the text analysis unit 13, and the text analysis unit 13 is instructed to execute the answer output process. This speech recognition result was acquired in step ST103 and, based on the evaluation result in step ST104, was determined to be appropriate as the subject of the answer in ST105. The question prompt corresponding to this speech recognition result is the same as the question prompt (after conversion) shown in Figure 4.

[0091] Furthermore, in step ST106, responses (questionnaire results), such as those shown in Figure 8, are output from the text analysis unit 13 to the response control unit 43. The output responses correspond to each question in the question prompt.

[0092] Next, the alert determination unit 55 determines whether an alert is necessary based on the response obtained in step ST106 (ST107). If it is determined that an alert is not necessary (No in ST107), the response control unit 43 stores the information of the obtained response in the response storage unit 45 (ST108).

[0093] In step ST108, information from multiple responses is stored in the response storage unit 45, as shown in Figure 9, for example. The response information includes identification information (in this case, a contact ID) for identifying the call between the customer and the operator, along with answers to each associated question (in this case, reasons for customer dissatisfaction regarding the operator's response, the operator's language, the operator's understanding, whether they will use the service again, the operator's response, etc., and customer opinions and feedback).

[0094] Subsequently, the response aggregation unit 47 retrieves the aggregated responses from the response information stored in the response storage unit 45 (ST109). The retrieved aggregated response information is stored in the aggregated response storage unit 49.

[0095] In step ST109, for example, the response aggregation prompt shown in Figure 10 is input from the response control unit 43 to the text analysis unit 13, and the text analysis unit 13 is instructed to execute the response aggregation process.

[0096] The response aggregation prompt includes information such as the respondent's role (in this case, a hypothetical professional editor), the type of text being asked about (in this case, customer opinions and feedback collected through a survey), and specific instructions (to aggregate similar opinions from the survey results). The response aggregation prompt also includes instructions regarding the input format (in this case, specific opinions and feedback) and the output format (in this case, the aggregated similar opinions and feedback).

[0097] In step ST109, along with the aforementioned prompt for aggregating responses, some of the response information stored in the response storage unit 45 (those with free-response answers among the survey items), such as the contact ID shown in Figure 11 and specific opinions and feedback from multiple customers (text to be standardized for inconsistencies in notation), is input from the response control unit 43 to the text analysis unit 13.

[0098] Furthermore, in step ST109, the text analysis unit 13 outputs the result of unifying variations in notation in the text (in this case, a collection of similar sentences related to opinions and impressions), as shown in Figure 12, to the response aggregation unit 47. In this result of unifying variations in notation, the information on the aggregated (i.e., unified) opinions and impressions is shown in association with the identification information of multiple corresponding calls (in this case, contact IDs).

[0099] In step ST109, as shown in Figure 13, for example, information of multiple aggregated responses is stored in the aggregated response storage unit 49 by combining the results of unifying variations in notation output from the text analysis unit 13 to the response aggregation unit 47 and the response information stored in the response storage unit 45, based on the identification information of multiple calls (in this case, contact ID). The aggregated responses include the content of the response information shown in Figure 9, as well as aggregated opinions and impressions based on the results of unifying variations in notation. This aggregation process facilitates statistical processing in the data analysis unit 53.

[0100] Step ST109 does not always need to be executed when responses based on the response aggregation process are obtained from the text analysis unit 13 (ST106), and may be omitted in some cases. In particular, if the questionnaire does not include free-response items such as specific opinions and impressions from customers, omitting the aggregation process is effective because there is no need to correct variations in notation.

[0101] Subsequently, the data analysis unit 53 performs data analysis, such as statistical processing, on the aggregated response data based on instructions for data analysis from the analyst, which are input from the analysis display unit 15 through the analysis instruction input unit 51 (ST110). For example, if an instruction is given to create a pie chart showing the number of aggregated opinions and comments for the aggregated responses in Figure 13, the unit will count and aggregate the same number of aggregated opinions and comments to create a pie chart, as shown in Figure 14, and display it on the analysis display unit 15. The data analysis in step ST110 can be started when an instruction for data analysis is input to the analysis instruction input unit 51.

[0102] On the other hand, if it is determined in step ST107 above that an alert is necessary (Yes), the alert unit 59 outputs (or transmits) alert information to the alert receiving unit 17 (ST111).

[0103] In step ST107, the responses of the response control unit 43 that determine that an alert is necessary include, for example, those in the survey results for the call center response items that match "dissatisfied," as shown in Figure 15.

[0104] In step ST111, the voice recognition result of the survey response that was determined to require an alert, as shown in Figure 16, and the summary prompt, as shown in Figure 17, are input from the alert summary control unit 57 to the text analysis unit 13.

[0105] Furthermore, in step ST111, for example as shown in Figure 18, a response (summary result) corresponding to the summarization prompt is output from the text analysis unit 13 to the alert summarization control unit 57.

[0106] Furthermore, in step ST111, as shown in Figure 19, for example, an alert notification containing the contact ID, call date and time, and operator name obtained from the call unit, as well as the summary results and survey results output from the text analysis unit 13 to the alert summary control unit 57, is output from the alert unit 59 to the alert receiving unit 17. This notifies the troubleshooter immediately after the call about any phone interaction that may have caused dissatisfaction with the customer. As a result, the troubleshooter can immediately follow up with the customer, preventing the situation from escalating into a serious problem.

[0107] Here, the alert notification includes information such as the date and time of the call, call identification information (in this case, contact ID), operator identification information (e.g., operator's name), customer identification information (e.g., the name of the store the customer belongs to and the customer's name), the content of the operator's response, the customer's evaluation of the operator's response, the customer's evaluation of the operator's language use, the customer's evaluation of the operator's level of understanding, the customer's evaluation of the problem resolution, whether the customer is likely to use the service again (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction, and the customer's opinions and feedback.

[0108] Furthermore, the recipient of alert notifications may be changed depending on the content of the survey responses. If a customer is dissatisfied with the service, notifying the troubleshooter allows for immediate follow-up with the customer. If the customer is angry, notifying not only the troubleshooter but also the operator's manager allows the manager to immediately follow up with the operator who handled the call, alleviating the operator's anxiety. Additionally, if the customer expresses gratitude, notifying the operator who handled the call can boost the operator's motivation.

[0109] Furthermore, if the person receiving the alert can easily check the relevant section of the audio when an alert is issued, they will be able to take more appropriate action, such as following up with the customer. Angry phone calls, in particular, can last for over 30 minutes, making it difficult to listen to the entire audio again for follow-up. By making it possible to easily play the problematic section of the audio when an alert is issued, the person receiving the alert can easily identify the problematic part of the call that triggered the alert.

[0110] In step ST103, the speech recognition control unit 37 obtains the speech recognition result of the call audio input to the voice input unit 35 from the speech recognition unit 9. This speech recognition result includes, for example, the time of each utterance by the customer and the operator, as shown in Figure 20.

[0111] In step ST106, the response control unit 43 inputs the speech recognition results determined to be appropriate (including the time of each utterance by the customer and the operator) as shown in Figure 20, and question prompts as shown in Figure 21, for example, into the text analysis unit 13. The question prompts include, for example, questions about the operator's response, the operator's wording, the operator's understanding, the resolution of any problems or questions the customer had, whether they would use the service again (i.e., whether the customer is likely to become a repeat customer), the reasons for the customer's dissatisfaction with the operator's response (an example of the basis for the answer), gratitude to the operator, the reason for the gratitude, anger towards the operator, the reason for the anger, questions to identify the problematic parts of the anger alert notification (start and end times of the relevant audio from the customer and the operator, the timing of the customer's anger, etc.), and questions regarding the customer's opinions and impressions.

[0112] In step ST106, in response to this question prompt, the text analysis unit 13 outputs an answer (including a description of the problem area), such as the one shown in Figure 24, to the answer control unit 43.

[0113] In step ST107, the alert determination unit 55 determines whether an alert is necessary and where to send the alert, based on the responses obtained in step ST106 and the rules shown in Figure 22 (including alert conditions and notification destinations). For example, if the survey response item "Anger" matches "Yes," an alert is necessary, and the notification destinations are determined to be "aaa@aaa.aaa.jp" and "bbb@bbb.bbb.jp".

[0114] In step ST111, for example as shown in Figure 23, an alert notification containing the contact ID, call date and time, and operator name obtained from the call unit, as well as the summary results and survey results output from the text analysis unit 13 to the alert summary control unit 57, is output from the alert unit 59 to the alert receiving unit 17, which is determined to be the notification destination.

[0115] Here, the alert notification includes information such as the date and time of the call, call identification information (in this case, contact ID), operator identification information (e.g., operator's name), customer identification information (e.g., the name of the store the customer belongs to and the customer's name), the content of the operator's response, the customer's evaluation of the operator's response, the customer's evaluation of the operator's language use, the customer's evaluation of the operator's level of understanding, the customer's evaluation of the problem resolution, whether the customer will use the service again (i.e., whether the customer is likely to become a repeat customer), the reason for the customer's dissatisfaction, whether the customer expressed gratitude for the operator's response and the reason for that gratitude, whether the customer was angry with the operator's response, the reason for that anger, the timing of that anger, and the customer's opinions and feedback.

[0116] The alert notification also includes a play button 60 for the problematic section. The troubleshooting staff can play back the problematic section of the call by pressing the play button 60 for the problematic section of the alert notification displayed on the alert receiving unit 17.

[0117] When the playback button 60 for the problematic section is pressed, the alert receiving unit 17 plays the corresponding call audio input to the audio input unit 35 from the start time to the end time. For example, if the answer is as shown in Figure 24, it will play from the start time: 00:00:31,000 to the end time: 00:01:03,000.

[0118] (Second Embodiment) Next, with reference to Figure 25 and other figures, the questionnaire processing system 1 according to the second embodiment will be described. Unlike questionnaires targeting people, this system allows for repeated questionnaires with modified questionnaire items. Therefore, by reviewing the results of the questionnaire before and after the modification with the questionnaire manager's eyes, it is possible to create a more effective questionnaire. In the drawings and descriptions relating to the second embodiment, components similar to those in the first embodiment described above are denoted by the same reference numerals as those used in the first embodiment. Furthermore, the questionnaire processing system 1 according to the second embodiment is the same as that of the first embodiment, except for matters specifically mentioned below.

[0119] In the questionnaire processing system 1 according to the second embodiment, it is possible to perform a process (hereinafter referred to as "questionnaire modification process") to modify the questionnaire items set by the questionnaire administrator to more appropriate content based on the answers output from the text analysis unit 13.

[0120] The survey processing system 1 includes an audio storage unit 63 that sequentially stores the audio of each call relayed by the call unit 5. The audio data of each call stored in the audio storage unit 63 is input to the audio input unit 35 as needed.

[0121] The main control unit 19 further includes a comparison information storage unit 65, a comparison result display unit 67, and a questionnaire determination unit 69.

[0122] The aggregated responses stored in the aggregated response storage unit 49 include a set of comparative responses (hereinafter referred to as the "comparative response set") obtained by question prompts based on different questionnaires (including questionnaire items that differ from each other) before and after the revision. The comparative information storage unit 65 stores (stores) the comparative response set extracted from the aggregated response storage unit 49.

[0123] The comparison results display unit 67 displays the comparison response sets stored in the comparison information storage unit 6 on the questionnaire setting unit 11.

[0124] In the survey setting unit 11, the survey administrator who has reviewed the comparison response set can modify the survey items as needed (i.e., select the modified survey items).

[0125] Information regarding the modification of the survey items is input from the survey setting unit 11 to the survey determination unit 69. Based on the information regarding the modification of the survey items, the survey determination unit 69 can input the modified survey items into the survey input unit 31.

[0126] As a result, the survey processing system 1 can have the text analysis unit 13 perform response output processing based on the revised survey items (question prompts) with more appropriate content.

[0127] Next, with reference to Figure 26, the flow of the questionnaire modification process performed in the questionnaire processing system 1 according to the second embodiment will be explained.

[0128] In the questionnaire correction process, ST201, which corresponds to step ST101 shown in Figure 3, is executed. In the subsequent step ST202, the audio file corresponding to each call stored in the audio storage unit 63 is acquired, and that audio file is input to the audio input unit 35. The speech recognition control unit 37 acquires the speech recognition result based on that audio file from the speech recognition unit 9 (ST203).

[0129] Next, steps ST203-ST205 and ST207-ST209, which correspond to steps STST103-ST106 and ST108-ST110 shown in Figure 3, are executed.

[0130] Subsequently, when the survey items are modified by the survey manager, the modified survey items are input from the survey setting unit 11 to the survey input unit 31 (ST210).

[0131] Therefore, the main control unit 19 performs processing based on the revised questionnaire (ST211). The processing in this step ST211 is the same as the processing in steps ST101, ST103-ST106, and ST108-ST110 shown in Figure 3. Furthermore, in step ST211, the comparison response set extracted from the aggregated response storage unit 49 is stored in the comparison information storage unit 65.

[0132] Next, the comparison result display unit 67 displays the comparison response set stored in the comparison information storage unit 65 on the questionnaire setting unit 11 (ST212). The comparison response set displayed there includes, for example, the questionnaire items before and after the change, and the corresponding responses (questionnaire results), as shown in Figure 27.

[0133] Next, when the survey administrator selects the revised survey items (i.e., determines that the revised survey items are more appropriate), information about the revisions to the survey items is entered from the survey setting unit 11 to the survey determination unit 69 (ST213). This determines the content of the new survey (survey items).

[0134] Subsequently, the questionnaire processing system 1 according to the second embodiment can perform the same processing as shown in Figure 3 above, based on the content of the newly determined questionnaire.

[0135] (Third embodiment) Next, with reference to Figure 28 and the like, the questionnaire processing system 1 according to the third embodiment will be described. In the first embodiment, speech recognition results with poor accuracy were excluded from the questionnaire. However, there are cases where the number of calls is small and it is not desirable to reduce the number of questionnaire subjects. Therefore, by correcting the speech recognition results in the text analysis unit, the accuracy of speech recognition can be improved, and the number of questionnaire subjects can be increased while maintaining the quality of the questionnaire responses. In the drawings and descriptions relating to the third embodiment, components similar to those in the first or second embodiment described above are denoted by the same reference numerals as those used in the first or second embodiment. Furthermore, the questionnaire processing system 1 according to the third embodiment is the same as in the first or second embodiment, except for matters specifically mentioned below.

[0136] In the questionnaire processing system 1 according to the third embodiment, the main control unit 19 further includes a text correction unit 71.

[0137] Here, the text analysis unit 13 can perform a text correction process (hereinafter referred to as "correction process") on the speech recognition results of the speech recognition unit 9. The text correction unit 71 instructs the text analysis unit 13 to perform the correction process on the speech recognition results obtained by the speech recognition control unit 37. As a result, the text correction unit 71 obtains the corrected speech recognition results from the text analysis unit 13.

[0138] The text correction unit 71 can input prompts to the text analysis unit 13, such as those shown in Figure 29 (hereinafter referred to as "correction prompts"), when instructing the text analysis unit 13 to perform correction processing.

[0139] Furthermore, the text analysis unit 13 can output the corrected speech recognition result (after correction) to the text correction unit 71 by performing a correction process on the speech recognition result (before correction), as shown in Figure 30, for example.

[0140] The speech recognition evaluation control unit 39 can instruct the text analysis unit 13 to execute speech recognition evaluation processing on the corrected speech recognition results obtained by the text correction unit 71.

[0141] (Fourth Embodiment) Next, with reference to Figure 31, the questionnaire processing system 1 according to the fourth embodiment will be described. If the system is configured to allow the operator to put calls on hold, and the call audio includes hold music, the accuracy of speech recognition during and before / after the hold period will be reduced. Therefore, by removing the hold music, the accuracy of speech recognition can be improved, and the quality of questionnaire responses can be increased. In the drawings and descriptions relating to the fourth embodiment, components that are the same as those in any of the first to third embodiments described above are denoted by the same reference numerals as those used in any of the first to third embodiments. Furthermore, the questionnaire processing system 1 according to the fourth embodiment is the same as in any of the first to third embodiments, except for matters specifically mentioned below.

[0142] In the questionnaire processing system 1 according to the fourth embodiment, the main control unit 19 further includes a hold sound removal unit 74.

[0143] The voice receiving unit 7 receives (collects) the audio of the conversation between the customer and the operator that takes place in the call unit 5, as well as the start and end of the operator's hold.

[0144] The hold tone removal unit 74 is located between the voice receiving unit 7 and the voice input unit 35 and receives call audio input from the voice receiving unit 7 and the operator's call hold start and end information. If the call audio contains telephone hold tone, the hold tone removal unit 74 uses the operator's call hold start and end information to remove the data related to the hold tone from the call audio data. As a result, the voice input unit 35 receives call audio from the hold tone removal unit 74 with the hold tone removed.

[0145] With the above configuration, the questionnaire processing system 1 according to the fourth embodiment eliminates the hold music included in the call audio, thus preventing the hold music from negatively affecting the responses output from the text analysis unit 13.

[0146] (Fifth embodiment) Next, with reference to Figure 32 and the like, the questionnaire processing system 1 according to the fifth embodiment will be described. When the text analysis unit analyzes the voice text of a call, it may misidentify the speaker. For example, when an operator is expressing gratitude to a customer, the system may misidentify that the customer is expressing gratitude to the operator. Misidentification can be prevented by inputting the voice recognition results, including speaker information, to the text analysis unit. In the drawings and descriptions relating to the fifth embodiment, components that are the same as those in any of the first to fourth embodiments described above are denoted by the same reference numerals as those used in any of the first to fourth embodiments. Furthermore, the questionnaire processing system 1 according to the fifth embodiment is the same as in any of the first to fourth embodiments, except for matters specifically mentioned below.

[0147] In the questionnaire processing system 1 according to the fifth embodiment, the main control unit 19 further includes a recognition result synthesis unit 75.

[0148] In the survey processing system 1, the audio data of the customer-operator conversation acquired by the audio receiving unit 7 (including customer audio data related to the customer's voice and operator audio data related to the operator's voice) is separated into L channel and R channel respectively and input to the audio input unit 35.

[0149] The speech recognition control unit 37 instructs the speech recognition unit 9 to perform speech recognition processing on the L channel and R channel audio input to the speech input unit 35, with each channel separated. As a result, the speech recognition control unit 37 obtains speech recognition results for the L channel and R channel audio, respectively.

[0150] The recognition result synthesis unit 75 synthesizes the speech recognition results for the L channel and R channel audio acquired by the speech recognition control unit 37. In this case, the synthesized speech recognition result may include, for example, the customer's voice and the operator's voice in an identifiable manner, as shown in Figure 33.

[0151] With the above configuration, the questionnaire processing system 1 according to the fifth embodiment can obtain appropriate answers to questionnaire items from the text analysis unit 13 based on call audio data (i.e., speech recognition results) in which the speakers (in this case, the customer and the operator) are distinguished.

[0152] (Sixth Embodiment) Next, with reference to Figure 34, the questionnaire processing system 1 according to the sixth embodiment will be described. In the drawings and description relating to the sixth embodiment, components similar to those in any of the first to fifth embodiments described above are denoted by the same reference numerals as those used in any of the first to fifth embodiments. Furthermore, the questionnaire processing system 1 according to the sixth embodiment is the same as that of any of the first to fifth embodiments, except for matters specifically mentioned below.

[0153] The questionnaire processing system 1 according to the sixth embodiment further includes a speech recognition unit 9B that employs a different speech recognition model (for example, a different machine learning model than 9A) in addition to the speech recognition unit 9A shown in Figure 1, which corresponds to the speech recognition unit 9A. The questionnaire processing system 1 may include three or more speech recognition units.

[0154] The speech recognition control unit 37 instructs multiple speech recognition units (in this case, speech recognition unit 9A and speech recognition unit 9B) to perform speech recognition processing for the audio of the call between the customer and the operator. As a result, the speech recognition control unit 37 can obtain speech recognition results from speech recognition unit 9A and speech recognition unit 9B, respectively.

[0155] The speech recognition evaluation control unit 39 can compare the speech recognition results obtained from the speech recognition unit 9A and the speech recognition unit 9B, respectively. For example, the speech recognition evaluation control unit 39 can compare multiple (in this case, two) speech recognition results using WER (Word Error Rate) and evaluate the differences (i.e., determine whether the multiple speech recognition results match or not).

[0156] The response target selection unit 41 determines whether or not to instruct the text analysis unit 13 to execute response output processing for the speech recognition result being evaluated by the speech recognition evaluation control unit 39, based on the evaluation result of the difference in speech recognition results. For example, the response target selection unit 41 can determine whether or not to instruct the response output processing to execute by comparing the WER value with a preset threshold.

[0157] With the above configuration, the questionnaire processing system 1 according to the sixth embodiment obtains answers to questionnaire items from the text analysis unit 13 only when there is a high probability that the speech recognition results are appropriate (i.e., when the speech recognition results obtained from multiple speech recognition units match), thus enabling the stable acquisition of appropriate answers to questionnaire items.

[0158] (Seventh Embodiment) Next, with reference to Figure 35, the questionnaire processing system 1 according to the seventh embodiment will be described. In the drawings and description relating to the seventh embodiment, components similar to those in any of the first to sixth embodiments described above are denoted by the same reference numerals as those used in any of the first to sixth embodiments. Furthermore, the questionnaire processing system 1 according to the seventh embodiment is the same as that of any of the first to sixth embodiments, except for matters specifically mentioned below.

[0159] The questionnaire processing system 1 according to the seventh embodiment further includes a call terminal control unit 77. The main control unit 19 further includes an ACW (After Call Work) control unit 79.

[0160] The call terminal control unit 77 controls which operator to assign to a customer inquiry (i.e., an incoming call from telephone 20) via the call unit 5 (i.e., which operator terminal 2 to call). At this time, an operator who is already on a call with another customer should not be called. Also, other tasks may arise after the call ends, such as taking notes about the content of the conversation. Furthermore, customer service can be mentally exhausting, and operators need time to compose themselves for the next call. For this reason, an operator who has just finished a call should not be called either. For this reason, it is common practice to use an ACW time (a period of time after a call has ended) to prevent calling operators who are currently on a call or have just finished a call. However, the time actually needed after a call ends varies depending on the content of the call. For example, if the customer the operator spoke to was angry, the operator will experience significant stress, and it will take longer than usual to compose themselves for the next call. Therefore, it is desirable to change the ACW time according to the content of the call.

[0161] Based on the call time and call duration (length of the call) information obtained from the response control unit 43, and the ACW time (downtime after the end of one call) set for each operator, the ACW control unit 79 can instruct the call terminal control unit 77 on the time until the operator starts a call with the next customer.

[0162] For example, as shown in Figure 24, the ACW control unit 79 can change the ACW time of the operator handling the call to be greater than the standard value (i.e., to suspend further calls for a predetermined time) if the response item of the response control unit for anger is "yes". This reduces the psychological anxiety of the operator who handled the call that triggered the alert by giving them a rest.

[0163] The first to seventh embodiments described an example in which the survey processing system 1 is applied to a survey of operators who handle customer inquiries over the phone at a call center. However, the survey processing system 1 may also be applied to a survey of teachers in a one-to-one online lesson, as described below.

[0164] (Eighth embodiment) Next, the questionnaire processing system 1 according to the eighth embodiment will be described with reference to Figure 36, etc. In the drawings and description relating to the eighth embodiment, components similar to those in any of the first to seventh embodiments described above are denoted by the same reference numerals as those used in any of the first to seventh embodiments. Furthermore, the questionnaire processing system 1 according to the eighth embodiment is the same as that of any of the first to seventh embodiments, except for matters specifically mentioned below.

[0165] The questionnaire processing system 1 according to the eighth embodiment has the same configuration as the questionnaire processing system 1 shown in Figure 1. However, in the eighth embodiment, instead of the operator terminal 2 and telephone 20 shown in Figure 1, a teacher terminal used by teachers and a student terminal used by students are used, respectively. In other words, in the eighth embodiment, the target of processing is not telephone calls related to customer inquiries in a call center as described above, but rather audio from online lessons between one teacher and one student (conversation between teacher and student), as shown in the speech recognition results of Figure 36, for example.

[0166] The prompt conversion unit 33 converts the survey items (before conversion) set by the survey administrator into question prompts (after conversion), as shown in Figure 37, for example.

[0167] Question prompts include information such as the respondent's role (in this case, a hypothetical elementary school boy taking online lessons) and the type of text being questioned (in this case, a transcript of a lesson). They also include questions about the teacher's response, language, understanding, resolution of student problems or questions, whether the student will use the service again (i.e., whether they will use online lessons again), reasons for student dissatisfaction with the teacher's response, and student opinions and feedback.

[0168] Furthermore, the text analysis unit 13 outputs, as a result of the response output processing, a response (questionnaire response) like the one shown in Figure 38 to the response control unit 43. This response includes answers to each question in the question prompt.

[0169] Furthermore, the questionnaire processing system 1 may be applied to questionnaires concerning doctors in online consultations between one doctor and one patient, as described below, for example.

[0170] (Ninth Embodiment) Next, the questionnaire processing system 1 according to the ninth embodiment will be described with reference to Figure 39, etc. In the drawings and description relating to the ninth embodiment, components similar to those in any of the first to eighth embodiments described above are denoted by the same reference numerals as those used in any of the first to eighth embodiments. Furthermore, the questionnaire processing system 1 according to the ninth embodiment is the same as in any of the first to eighth embodiments, except for matters specifically mentioned below.

[0171] The questionnaire processing system 1 according to the eighth embodiment has the same configuration as the questionnaire processing system 1 shown in Figure 1. However, in the ninth embodiment, instead of the operator terminal 2 and telephone 20 shown in Figure 1, a doctor's terminal used by a doctor and a patient's terminal used by a patient are used, respectively. In other words, in the ninth embodiment, the target of processing is not telephone calls from customers in a call center as described above, but rather audio from online consultations between one doctor and one patient (conversations between a doctor and a patient), as shown in the speech recognition results of Figure 39, for example.

[0172] The prompt conversion unit 33 converts the survey items (before conversion) set by the survey administrator into question prompts (after conversion), as shown in Figure 40, for example.

[0173] Question prompts include information such as the respondent's role (in this case, a hypothetical patient receiving online medical care) and the type of text being questioned (in this case, a transcript of the medical consultation). They also include questions about the doctor's response, language, understanding, resolution of patient problems and questions, whether the patient will use the service again (i.e., whether they will use online medical care again), reasons for patient dissatisfaction with the doctor's response, and patient opinions and feedback.

[0174] Furthermore, the text analysis unit 13 outputs, as a result of the response output processing, a response (questionnaire response) like the one shown in Figure 41 to the response control unit 43. This response includes answers to each question in the question prompt.

[0175] As described above, embodiments have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these embodiments and can be applied to embodiments that have been modified, replaced, added, or omitted. Furthermore, it is possible to create new embodiments by combining the components described in the above embodiments.

[0176] For example, in the above-described embodiment, an example was shown in which the survey processing system 1 is applied to a telephone call or a similar communication method. However, the scope of the survey processing system 1 is not limited to telephone calls. For example, in the survey processing system 1, speech recognition (processing that converts speech to text), etc., may be omitted, and survey processing may be performed on a conversation that has already been transcribed into text (for example, chat text). [Industrial applicability]

[0177] The questionnaire processing system, questionnaire processing method, and questionnaire processing program relating to this disclosure have the effect of being able to obtain the results of questionnaires regarding the responses of operators who have spoken with customers, with a simple configuration, and are useful as a questionnaire processing system, questionnaire processing method, and questionnaire processing program for obtaining questionnaire results. [Explanation of Symbols]

[0178] 1: Questionnaire processing system 2: Operator terminal 3: Operator terminal 5:Talking part 6: Accumulation of comparative information 7: Audio receiving unit 9: Speech Recognition Unit 11: Survey Setup Department 13: Text Analysis Department 15:Analysis display section 17: Alert receiving unit 19: Main Control Unit 20: Telephone 21: Display device 22: Input Devices 23: Mike 24: Speaker 31: Questionnaire Input Section 33: Prompt conversion section 35: Voice input section 37: Speech Recognition Control Unit 39: Speech Recognition Evaluation Control Unit 41: Selection Department for Respondents 43: Response Control Unit 45: Answer storage section 47: Answer Compilation Department 49: Aggregated Response Storage Department 51: Analysis instruction input section 53: Data Analysis Department 55: Alert Judgment Unit 57: Alert Summarization Control Unit 59: Alert Department 60: Play button 63: Audio Storage Unit 65: Accumulation of comparative information 67: Comparison result display section 69: Questionnaire Decision-Making Department 71: Text Correction Section 74: Hold music removal section 75: Recognition result synthesis section 77: Call Terminal Control Unit 79: ACW Control Unit

Claims

1. A survey processing system for obtaining survey responses from the customer's perspective, regarding a conversation between a customer and an operator, The system comprises one or more processors that perform processing to obtain the responses to the aforementioned questionnaire, The one or more processors mentioned above are: Obtain the survey items as questions related to the aforementioned survey, The audio of the aforementioned conversation is acquired, By instructing a speech recognition unit capable of performing speech recognition to perform speech recognition on the conversational audio, the speech recognition unit obtains the conversational audio converted into text as a result of speech recognition. An answer processing system that obtains answers to the survey items from an answer output unit, which is capable of outputting answers to questions in a transcribed conversational audio, by instructing the answer output unit to output answers to the survey items in the speech recognition results.

2. The one or more processors mentioned above are: By instructing a speech recognition evaluation unit, which is capable of evaluating the results of speech recognition, to evaluate the speech recognition results, the evaluation results are obtained from the speech recognition evaluation unit. The questionnaire processing system according to claim 1, which determines whether or not to instruct the response output unit to output the response based on the evaluation result.

3. The aforementioned questionnaire items are set based on human input, The aforementioned response output unit is composed of a large-scale language model, The one or more processors mentioned above are: The questionnaire processing system according to claim 1, comprising converting the questionnaire items into prompts and instructing the response output unit to output responses to the questionnaire items in the speech recognition results based on the prompts.

4. The response obtained from the response output unit includes the customer's opinions and impressions. The one or more processors mentioned above are: The answers obtained from the answer output unit are sequentially stored, The questionnaire processing system according to claim 1, wherein the system instructs the response output unit to unify inconsistencies in notation in the accumulated multiple response sentences, thereby obtaining multiple responses from the response output unit that include sentences with unified notation.

5. The one or more processors mentioned above are: By instructing multiple speech recognition units capable of evaluating the results of speech recognition to perform speech recognition on the conversational audio, the speech recognition results are obtained from each of the multiple speech recognition units. Determine whether the multiple speech recognition results match or not. The questionnaire processing system according to claim 1, wherein the system instructs the response output unit to output the response only when it is determined that the plurality of speech recognition results match.

6. The one or more processors mentioned above are: Based on the responses to the survey items obtained from the response output unit, it is determined whether an alert is necessary. The questionnaire processing system according to claim 1, which, when it is determined that the aforementioned alert is necessary, outputs an alert regarding the answers to the questionnaire items to a pre-configured alert receiving unit.

7. The aforementioned conversation is an online conversation between a first user and a second user. The first user conducts a conversation with the second user using the first user terminal. The one or more processors mentioned above are: The questionnaire processing system according to claim 6, wherein, when the aforementioned alert is output, the system instructs the terminal control unit that controls calls to the first user terminal to suspend further calls for a predetermined period of time.

8. The one or more processors mentioned above are: The questionnaire processing system according to claim 1, wherein when instructing the response output unit to output a response to the questionnaire item in the speech recognition result, the system also instructs the output of the basis for the response.

9. The one or more processors mentioned above are: If the aforementioned conversation is an online conversation and includes hold music, remove the hold music. The questionnaire processing system according to claim 1, wherein the voice recognition unit is instructed to perform voice recognition on the conversational audio from which the data relating to the hold music has been removed.

10. The one or more processors mentioned above are: The first audio, related to the voice of the first user, and the second audio, related to the voice of the second user, are obtained, The first voice converted into text and the second voice converted into text are obtained as the voice recognition results from the voice recognition unit, The questionnaire processing system according to claim 1, wherein the voice recognition result for instructing the response output unit to output the responses to the questionnaire items includes, in an identifiable manner, the voice of the first user and the voice of the second user.

11. A survey processing method for obtaining survey responses from the customer's perspective as a participant in a conversation between a customer and an operator, One or more computers, Obtain the survey items as questions related to the aforementioned survey, The audio of the aforementioned conversation is acquired, By instructing a speech recognition unit capable of performing speech recognition to perform speech recognition on the conversational audio, the speech recognition unit obtains the conversational audio converted into text as a result of speech recognition. A survey processing method comprising obtaining answers to survey items from an answer output unit, which is capable of outputting answers to questions in a transcribed conversational audio, by instructing the answer output unit to output answers to the survey items in the speech recognition results.

12. A survey processing program that causes a computer to perform information processing to obtain survey responses from the customer's perspective, as a participant in a conversation between a customer and an operator, The aforementioned information processing includes, Obtain the survey items as questions related to the aforementioned survey, The audio of the aforementioned conversation is acquired, By instructing a speech recognition unit capable of performing speech recognition to perform speech recognition on the conversational audio, the speech recognition unit obtains the conversational audio converted into text as a result of speech recognition. A survey processing program that includes a procedure for obtaining answers to survey items from an answer output unit, which is capable of outputting answers to questions in transcribed conversational audio, by instructing the answer output unit to output answers to the survey items in the speech recognition results.

Citation Information

Patent Citations

  • System, program, and method for surveying

    JP7101357B2