Information processing system, information processing method and program
The system addresses the challenge of selecting the right speech recognition dictionary by dynamically changing it during calls, ensuring highly accurate speech-to-text conversion and enhancing response quality and analysis.
Patent Information
- Application Number
- JP2025051869
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-07-21
AI Technical Summary
Conventional speech recognition systems struggle with obtaining accurate results due to the difficulty in selecting the appropriate speech recognition dictionary, often defaulting to a general-purpose dictionary, leading to insufficient accuracy.
An information processing system that includes a selection unit to choose from multiple speech recognition dictionaries and performs speech recognition using the selected dictionary, with the option to change dictionaries during a call to ensure accurate conversion of speech to text.
This approach enables highly accurate speech recognition by allowing dynamic dictionary selection and reprocessing with appropriate dictionaries, improving response quality and analysis accuracy.
Smart Images

Figure 2025098176000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system, an information processing method, and a program.
Background Art
[0002] In voice recognition technology, generally, a voice recognition dictionary in which notations, readings, arrangements, etc. of words are registered is used. There are various types of such voice recognition dictionaries depending on the application, language, etc. for which voice recognition is targeted. For example, there are general-purpose dictionaries, dictionaries in which many technical terms related to specific operations are registered, dictionaries specialized for specific languages, dictionaries specialized for specific dialects, and the like.
[0003] In a contact center (or also called a call center), by a voice recognition system implementing the above voice recognition technology, the voice during a call is converted into text in real time, and the text is presented to an operator (for example, Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, conventionally, even when a plurality of speech recognition dictionaries were prepared, it was difficult for an operator to select an appropriate dictionary from among them. For this reason, speech recognition was performed using a speech recognition dictionary that was preset for the operator (for example, a general-purpose speech recognition dictionary set as the default), and as a result, there were cases where a speech recognition result with sufficient accuracy could not be obtained.
[0006] The present disclosure has been made in view of the above points, and an object thereof is to provide a technology capable of obtaining a highly accurate speech recognition result.
Means for Solving the Problems
[0007] An information processing system according to an aspect of the present disclosure includes a selection unit that selects a speech recognition dictionary to be used for speech recognition from among a plurality of speech recognition dictionaries, and a speech recognition unit that uses the speech recognition dictionary selected by the selection unit to generate a speech recognition text obtained by converting speech included in a voice call with a customer into text by the speech recognition. When the speech recognition dictionary selected by the selection unit is changed, a speech recognition text obtained by converting some or all of the speech included in the voice call into text by the speech recognition in parallel using the changed speech recognition dictionary is generated.
Effects of the Invention
[0008] A technology capable of obtaining a highly accurate speech recognition result is provided.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0010] Hereinafter, an embodiment of the present invention will be described. Hereinafter, in this embodiment, for a contact center, when it is possible to automatically or manually select a dictionary from a plurality of voice recognition dictionaries, a contact center system 1 capable of obtaining an accurate voice recognition result for the voice of a call between an operator and a customer will be described. However, the contact center is an example, and for example, when it is possible to automatically or manually select a dictionary from a plurality of voice recognition dictionaries for an office or the like, it can be similarly applied when obtaining an accurate voice recognition result for the voice of a call between a person in charge and a customer.
[0011] <Overall Configuration of Contact Center System 1> Fig. 1 shows an example of the overall configuration of the contact center system 1 according to this embodiment. As shown in Fig. 1, the contact center system 1 according to this embodiment includes a voice recognition system 10, a plurality of user terminals 20, a plurality of telephones 30, a PBX (Private Branch eXchange) 40, an NW switch 50, and a customer terminal 60. Here, the voice recognition system 10, the user terminals 20, the telephones 30, the PBX 40, and the NW switch 50 are installed in a contact center environment E which is a system environment of the contact center. Note that the contact center environment E is not limited to a system environment within the same building, and may be, for example, a system environment within a plurality of geographically separated buildings.
[0012] The voice recognition system 10 uses the packets (voice packets) transmitted from the NW switch 50 to record the voice of the call between the operator and the customer as a voice file. Note that the voice recognition system 10 may passively acquire the voice packets transmitted from the NW switch 50, or may actively acquire the voice data by requesting the PBX 40 for voice data via the NW switch 50.
[0013] Further, the voice recognition system 10 performs voice recognition on this voice file and generates a text representing the voice recognition result (hereinafter also referred to as a voice recognition text). At this time, when the voice recognition dictionary is changed, the voice recognition system 10 performs voice recognition again on the voice files that have already been voice recognized using the changed voice recognition dictionary (that is, performs voice recognition including the voices that have already been voice recognized using the previous voice recognition dictionary). Thereby, for example, when the voice recognition dictionary is changed from an inappropriate one to an appropriate one, the voices that have already been voice recognized by the inappropriate voice recognition dictionary are voice recognized again by the appropriate voice recognition dictionary, and it becomes possible to obtain a highly accurate voice recognition result. Note that the voice recognition system 10 is realized by, for example, a general-purpose server or a server group.
[0014] The user terminal 20 is a terminal such as a PC (Personal Computer) used by a user (operator or supervisor). In the following, the user is mainly assumed to be an operator, but some users may be supervisors. Note that an operator is a person whose main job is to answer calls from customers, etc. On the other hand, a supervisor is a person who monitors the calls of operators and supports the telephone answering operations of those operators when some problem is likely to occur or in response to requests from the operators. Usually, the calls of several to a dozen or so operators are generally monitored by one supervisor.
[0015] On the user terminal 20, a response support screen is displayed in which the voice recognition result (voice recognition text) during a call with a customer is visualized in real time. By referring to this response support screen, the operator can also confirm the content of the call with the customer as text.
[0016] The telephone 30 is an IP (Internet Protocol) telephone (such as a fixed IP telephone or a mobile IP telephone) used by the operator.
[0017] The PBX 40 is a telephone exchange (IP-PBX) and is connected to a communication network 70 including a VoIP (Voice over Internet Protocol) network and a PSTN (Public Switched Telephone Network).
[0018] The NW switch 50 relays packets between the telephone 30 and the PBX 40 and captures those packets and sends them to the voice recognition system 10.
[0019] The customer terminal 60 is various terminals such as a smartphone, a mobile phone, or a fixed phone used by the customer.
[0020] Note that the overall configuration of the contact center system 1 shown in FIG. 1 is an example, and other configurations may be used. For example, in the example shown in FIG. 1, the speech recognition system 10 is included in the contact center environment E (that is, the speech recognition system 10 is on-premises), but all or some of the functions of the speech recognition system 10 may be realized by a cloud service or the like. Similarly, in the example shown in FIG. 1, the PBX 40 is an on-premises telephone switch, but it may be realized by a cloud service. Further, when the user terminal 20 has a telephone function, the contact center system 1 may not include the telephone 30.
[0021] <Functional Configuration of Contact Center System 1> FIG. 2 shows an example of the functional configurations of the speech recognition system 10 and the user terminal 20 included in the contact center system 1 according to the present embodiment.
[0022] ≪Speech Recognition System 10≫ As shown in FIG. 2, the speech recognition system 10 according to the present embodiment includes a speech recording unit 101, a dictionary selection unit 102, a speech recognition unit 103, and a UI providing unit 104. Each of these units is realized, for example, by processing in which one or more programs installed in the speech recognition system 10 are executed by a processor such as a CPU (Central Processing Unit). Further, the speech recognition system 10 according to the present embodiment includes a speech storage unit 105, a dictionary storage unit 106, and a call history storage unit 107. Each of these units can be realized by a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. However, at least a part of the storage areas of these units may be realized by a storage device (such as a database server) that is communicably connected to the speech recognition system 10.
[0023] The speech recording unit 101 stores the speech data represented by the packet (speech packet) transmitted from the NW switch 50 as a speech file in the speech storage unit 105.
[0024] The dictionary selection unit 102 selects a speech recognition dictionary 500 to be used for speech recognition from among a plurality of speech recognition dictionaries 500 stored in the dictionary storage unit 106. The speech recognition dictionary 500 is, for example, dictionary information in which the notation of words, their pronunciations, word order, etc. are registered. Examples of the speech recognition dictionary 500 include speech recognition dictionaries for general-purpose use, speech recognition dictionaries specialized for specific operations (e.g., finance, insurance, information and communication, etc.), speech recognition dictionaries specialized for specific languages (e.g., Japanese, English, French, etc.), speech recognition dictionaries specialized for specific dialects (e.g., dialects of a certain region in Japan, etc.), and various other types of speech recognition dictionaries. Hereinafter, the speech recognition dictionary 500 selected by the dictionary selection unit 102 will also be referred to as the "selected dictionary 500".
[0025] The speech recognition unit 103 performs speech recognition on the speech files stored in the speech storage unit 105 using the selected dictionary 500 selected by the dictionary selection unit 102, and generates a speech recognition text that is the result of the speech recognition. At this time, the speech recognition unit 103 performs speech recognition for each speaker (operator, customer), and generates a speech recognition text with speaker information and time information. The speech recognition text for a certain sentence (one utterance, one phrase, etc.) is represented in a form such as (speaker information, time information, speech recognition text), for example. Such a speech recognition text with speaker information and time information can be generated by known speech recognition techniques. Note that the speaker information is information indicating the speaker (operator or customer) who uttered the speech corresponding to the speech recognition text, and the time information is information indicating the time (date and time) when the speech corresponding to the speech recognition text was uttered. Hereinafter, it is assumed that the speech recognition text is provided with speaker information and time information, and is represented in a form such as (speaker information, time information, speech recognition text).
[0026] Also, when the selected dictionary 500 is changed, the speech recognition unit 103 performs speech recognition again on the speech files that have already been speech-recognized using the changed selected dictionary 500.
[0027] Furthermore, when a call between an operator and a customer ends, for example, the voice recognition unit 103 stores call history information including voice recognition text related to the call in the call history storage unit 107.
[0028] The UI providing unit 104 provides screen information of a response support screen on which the voice recognition text generated by the voice recognition unit 103 is visualized. Note that the screen information is represented by information such as, for example, HTML (Hypertext Markup Language), CSS (Cascading Style Sheets), JavaScript, etc.
[0029] The voice storage unit 105 stores a voice file of the voice represented by the packet (voice packet) transmitted from the NW switch 50.
[0030] The dictionary storage unit 106 stores a plurality of voice recognition dictionaries 500. Assume that among these plurality of voice recognition dictionaries 500, there is a voice recognition dictionary 500 (hereinafter referred to as the "default dictionary 500") selected as the default (standard). The default dictionary 500 is generally often a voice recognition dictionary for general-purpose use. However, for example, when mainly handling inquiries for specific operations at a contact center, a voice recognition dictionary specialized for that operation may be set as the default dictionary 500. Or, for example, when mainly handling inquiries for customers in a specific language at a contact center, a voice recognition dictionary specialized for that language may be set as the default dictionary 500, or when handling inquiries for customers in a specific region, a voice recognition dictionary specialized for the dialect of that region may be set as the default dictionary 500.
[0031] The call history storage unit 107 stores call history information. The call history information is information that includes, for example, at least a call ID and voice recognition text related to the call with that call ID. Note that the call history information may include various information such as, for example, the call date and time, call duration, the ID of the operator who handled the call, the extension number of the operator, the customer's phone number, and any memo information related to the call.
[0032] ≪User Terminal 20≫ As shown in FIG. 2, the user terminal 20 according to the present embodiment has a UI control unit 201. The UI control unit 201 is realized, for example, by a process that causes a processor such as a CPU to execute one or more programs (such as a web browser) installed in the user terminal 20.
[0033] The UI control unit 201 displays various screens including a response support screen on the display of the user terminal 20. Further, the UI control unit 201 receives various input operations of the user on these various screens.
[0034] <Response Support Process> Hereinafter, a process (response support process) of performing voice recognition on the voice of a call between an operator and a customer during the call and displaying the voice recognition result on the response support screen of the user terminal 20 will be described with reference to FIG. 3.
[0035] When a call between an operator and a customer is started, the voice recording unit 101 of the voice recognition system 10 receives a packet (start packet) indicating that the call has started (step S101).
[0036] Next, the dictionary selection unit 102 of the voice recognition system 10 selects a voice recognition dictionary 500 to be used for voice recognition from among a plurality of voice recognition dictionaries 500 stored in the dictionary storage unit 106 (step S102). Here, the dictionary selection unit 102 may, for example, select the default dictionary 500, or may inquire the user terminal 20 about which voice recognition dictionary 500 to use and then select the voice recognition dictionary 500 specified by the user (operator) in response to this inquiry. Also, when inquiring the user terminal 20 about which voice recognition dictionary 500 to use, the dictionary selection unit 102 may, for example, give the user (operator) a certain grace period of about several tens of seconds and select the default dictionary 500 when no specification of the voice recognition dictionary 500 is made within this grace period (in this case, voice recognition is not performed until the grace period has elapsed). This is because it is generally difficult for the operator to determine which voice recognition dictionary 500 should be used at the start of a call. Or, alternatively, for example, until a voice recognition dictionary 500 is explicitly selected by the operator, it may be regarded as if the default dictionary 500 has been selected.
[0037] The following steps S103 to S108 are repeatedly executed during the call between the operator and the customer.
[0038] The voice recording unit 101 of the voice recognition system 10 receives the packet (voice packet) transmitted from the NW switch 50 (step S103).
[0039] Next, the voice recording unit 101 of the voice recognition system 10 stores the voice data represented by the packet as a voice file in the voice storage unit 105 (step S104).
[0040] Next, the speech recognition unit 103 of the speech recognition system 10 performs speech recognition on the speech file stored in the speech memory unit 105 using the selected dictionary 500, and generates a speech recognition text that is the result of the speech recognition (step S105). At this time, if the selected dictionary 500 is changed in step S108 described later, the speech recognition unit 103 performs speech recognition again on the speech files that have already been speech recognized using the changed selected dictionary 500. Note that the details of the speech recognition in this step will be described later.
[0041] Next, the UI providing unit 104 of the speech recognition system 10 transmits the speech recognition text generated in step S105 above and the screen information for visualizing the speech recognition text to the user terminal 20 (for example, the user terminal 20 used by the operator making the call) (step S106). Here, the UI providing unit 104 may transmit the speech recognition text and the screen information to the user terminal 20 each time the speech recognition text is generated in step S105 above, or may transmit the speech recognition text and the screen information to the user terminal 20 in response to a request from the user terminal 20. Note that the UI providing unit 104 may transmit the speech recognition text and the screen information not only to the user terminal 20 used by the operator making the call, but also to, for example, the user terminal 20 used by a supervisor who monitors the call of the operator.
[0042] When the UI control unit 201 of the user terminal 20 receives the speech recognition text and the screen information, it displays the speech recognition text on the response support screen based on this screen information (step S107). Note that the details of the response support screen in this step will be described later.
[0043] When changing the selected dictionary 500, the dictionary selection unit 102 of the speech recognition system 10 changes the selected dictionary 500 to any one of the plurality of speech recognition dictionaries 500 (step S108). Here, for example, when a speech recognition dictionary 500 is specified by a user (operator), the dictionary selection unit 102 may change the selected dictionary 500 to that speech recognition dictionary 500. This is because after a certain amount of conversation has taken place, the operator can determine which speech recognition dictionary 500 should be used.
[0044] However, not limited to this, the dictionary selection unit 102 may determine whether to change the selected dictionary 500 by some judgment logic and determine which speech recognition dictionary 500 to change to. For example, the dictionary selection unit 102 may, after identifying the language in which the conversation is being conducted by known natural language processing, change the selected dictionary 500 to a speech recognition dictionary 500 specialized for the identified language. Similarly, for example, the dictionary selection unit 102 may, after identifying the customer's dialect by known natural language processing, change the selected dictionary 500 to a speech recognition dictionary 500 specialized for the identified dialect. Or, for example, the dictionary selection unit 102 may, based on the frequency of specific words and the like included in the previous speech recognition text (for example, the speech recognition result using the default dictionary 500 which is a speech recognition dictionary 500 for general purposes) by known inference techniques such as machine learning, infer the business content and then change the selected dictionary 500 to a speech recognition dictionary 500 specialized for that business.
[0045] When the call between the operator and the customer ends, the speech recognition unit 103 of the speech recognition system 10 creates call history information including the speech recognition text related to that call and stores the call history information in the call history storage unit 107 (step S109). Note that the call history information is used, for example, for various analyses to improve the response quality to the customer and the evaluation of the operator.
[0046] <Details of Speech Recognition in Step S105 of FIG. 3> The following describes the details of speech recognition in step S105 of FIG. 3. Hereinafter, it is assumed that the default dictionary 500 is selected in step S102 of FIG. 3.
[0047] ·Speech Recognition Example 1: When there is no change in the selected dictionary 500 As shown in FIG. 4, it is assumed that the speech recognition texts of utterances 1001 to 1008 are obtained by speech recognition using the default dictionary 500 at the time of the call time "00:35". Note that utterances 1001, 1003, 1005, and 1007 are the operator's utterances, and utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0048] At this time, in this speech recognition example, since there is no change in the selected dictionary 500, the speech recognition text of the operator's utterance 1009 at the call time "00:38" and the speech recognition text of the customer's utterance 1010 at the call time "00:43" are both obtained by speech recognition using the default dictionary 500.
[0049] Similarly, the speech recognition text of the operator's utterance 1011 at the call time "00:49" and the speech recognition text of the customer's utterance 1012 at the call time "00:54" are both obtained by speech recognition using the default dictionary 500.
[0050] Thus, when there is no change in the selected dictionary 500, the speech (utterance) during the call is speech recognized using the selected dictionary 500.
[0051] ·Speech Recognition Example 2: When the selected dictionary 500 is changed As shown in FIG. 5, it is assumed that the speech recognition texts of utterances 1001 to 1008 are obtained by speech recognition using the default dictionary 500 at the time of the call time "00:35". Note that utterances 1001, 1003, 1005, and 1007 are the operator's utterances, and utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0052] At this time, it is assumed that the selected dictionary 500 was changed after the call time of "00:35" and before the call time of "00:38". In this case, in this voice recognition example, using the changed selected dictionary 500, the already voice-recognized utterances 1001 to 1008 are voice-recognized in chronological order. On the other hand, regarding the utterances 1009 to 1012 after the change of the selected dictionary 500, after the voice recognition of the utterances 1001 to 1008 is completed, they are voice-recognized in chronological order.
[0053] In the example shown in FIG. 5, at the time of the call time of "00:45", the voice recognition text of voice recognition using the changed selected dictionary 500 for the utterances 1001 to 1003 is obtained. Also, at the time of the call time of "00:55", the voice recognition text of voice recognition using the changed selected dictionary 500 for the utterances 1001 to 1012 is obtained.
[0054] As described above, when the selected dictionary 500 is changed, in this voice recognition example, after re-voice-recognizing the utterances before the change in chronological order using the changed selected dictionary 500, the utterances after the change are voice-recognized in chronological order using the changed selected dictionary 500. Hereinafter, the utterances of the operator and the customer before the change of the selected dictionary 500 will also be referred to as "past utterances", and the utterances of the operator and the customer after the change of the selected dictionary 500 will also be referred to as "real-time utterances". Also, the voice file in which the voice of the past utterances is recorded will also be referred to as the "past voice file", and the voice file in which the voice of the real-time utterances is recorded will also be referred to as the "real-time voice file". Note that when the voice of the past utterances and the voice of the real-time utterances are recorded in the same voice file, the past voice file and the real-time voice file are the same voice file, but the voice of the past utterances and the voice of the real-time utterances may be recorded in different voice files. In this case, the past voice file and the real-time voice file are different voice files.
[0055] ·Voice Recognition Example 3: When the selected dictionary 500 is changed and the past voice files are processed in parallel for each utterance section In the above voice recognition example Part 2, the past utterances are recognized again in chronological order using the changed in-selection dictionary 500. This is because, generally, in voice recognition processing, it is necessary to perform voice recognition in order from the beginning of the voice file. On the other hand, by performing a process called voice activity detection (VAD) on the voice file, it becomes possible to perform voice recognition in parallel for each voice segment. Therefore, in this voice recognition example, after performing voice segment detection on the past voice file, the past utterances are recognized in parallel. However, the number of voices that can be recognized in parallel (hereinafter also referred to as the parallel number) depends on the number of voice recognition engines, etc., and is a predetermined number.
[0056] As shown in FIG. 6, assume that the voice recognition texts of utterances 1001 to 1008 at the time of the call time "00:35" are obtained by voice recognition using the default dictionary 500. Note that utterances 1001, 1003, 1005, and 1007 are the operator's utterances, and utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0057] At this time, assume that the in-selection dictionary 500 was changed after the call time "00:35" and before the call time "00:38". In this case, in this voice recognition example, using the changed in-selection dictionary 500, the already voice-recognized utterances 1001 to 1008 are recognized in parallel. On the other hand, regarding utterances 1009 to 1012 after the change of the in-selection dictionary 500, after the voice recognition of utterances 1001 to 1008 is completed, they are recognized in chronological order.
[0058] In the example shown in FIG. 6, at the time of the call time "00:45", the voice recognition texts of voice recognition using the changed in-selection dictionary 500 for utterances 1001 and 1004 to 1005 are obtained. This example is a case where the parallel number is 2, and utterances 1001 and 1004 to 1005 are recognized in parallel. Also, at the time of the call time "00:55", the voice recognition texts of voice recognition using the changed in-selection dictionary 500 for utterances 1001 to 1012 are obtained.
[0059] Thus, when the currently selected dictionary 500 is changed, in this speech recognition example, after the speech before the change is recognized again in parallel using the changed currently selected dictionary 500, the speech after the change is recognized in chronological order using the changed currently selected dictionary 500. Thereby, for example, speech recognition can be performed with priority for past speech. For example, it becomes possible to preferentially recognize speech that is closer to the real time among past speech and speech that is closer to the start of the call. Also, since past speech is recognized in parallel, it is possible to complete the speech recognition of past speech quickly.
[0060] Note that in this speech recognition example, speech recognition is performed in parallel for each speech segment by performing a process called speech segment detection, but this is just an example, and for example, sentences, phrases, etc. may be detected and speech recognition may be performed in parallel for each sentence unit, phrase unit, etc.
[0061] ·Speech Recognition Example 4: When the currently selected dictionary 500 is changed and past audio files and real-time audio files are processed in parallel In the above Speech Recognition Example 2, after all past speech is recognized using the changed currently selected dictionary 500, real-time speech is recognized using the changed currently selected dictionary 500. In contrast, by recording past speech and real-time speech in different audio files, it is possible to perform speech recognition on past speech and real-time speech in parallel. Therefore, in this speech recognition example, past speech and real-time speech are recorded in different audio files, and speech recognition is performed on past speech and real-time speech in parallel.
[0062] As shown in FIG. 7, assume that the speech recognition texts of speeches 1001 to 1008 are obtained by speech recognition using the default dictionary 500 at the time of the call time "00:35". Note that speeches 1001, 1003, 1005, and 1007 are the operator's speeches, and speeches 1002, 1004, 1006, and 1008 are the customer's speeches.
[0063] At this time, assume that the currently selected dictionary 500 was changed after the call time of "00:35" and before the call time of "00:38". In this case, in this voice recognition example, using the changed currently selected dictionary 500, the already voice-recognized utterances 1001 to 1008 are voice-recognized in chronological order, and the utterances 1009 to 1012 are also voice-recognized in chronological order. That is, the past utterances and the real-time utterances are voice-recognized in parallel and in chronological order.
[0064] In the example shown in FIG. 7, at the time of the call time "00:45", the voice recognition text of voice recognition using the changed currently selected dictionary 500 for the utterances 1001 to 1002 and the utterance 1009 is obtained. This example is a case where the past utterances 1001 to 1002 and the real-time utterance 1009 are voice-recognized in parallel. Also, at the time of the call time "00:55", the voice recognition text of voice recognition using the changed currently selected dictionary 500 for the utterances 1001 to 1012 is obtained.
[0065] As described above, when the currently selected dictionary 500 is changed, in this voice recognition example, using the changed currently selected dictionary 500, the utterances before the change and the utterances after the change are voice-recognized in parallel and in chronological order. Thereby, for example, it becomes possible to perform voice recognition of real-time utterances while simultaneously voice-recognizing past utterances.
[0066] ·Voice Recognition Example No. 5: When the currently selected dictionary 500 is changed, and the past voice files are processed in parallel in units of utterance sections, and the past voice files and the real-time voice files are processed in parallel This voice recognition example is a combination of the above voice recognition example No. 3 and voice recognition example No. 4. That is, in this voice recognition example, the past utterances and the real-time utterances are recorded in different voice files, and after detecting the utterance sections for the past voice files, the past utterances and the real-time utterances are voice-recognized in parallel, and the past utterances are also voice-recognized in parallel. However, the number of parallel processes for the past utterances depends on the number of voice recognition engines, etc., and is a predetermined number.
[0067] As shown in FIG. 8, it is assumed that the speech recognition texts of utterances 1001 to 1008 at the call time of "00:35" are obtained by speech recognition using the default dictionary 500. Note that utterances 1001, 1003, 1005, and 1007 are the operator's utterances, and utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0068] At this time, it is assumed that the selected dictionary 500 is changed after the call time of "00:35" and before the call time of "00:38". In this case, in this speech recognition example, using the changed selected dictionary 500, the already speech-recognized utterances 1001 to 1008 and utterances 1009 to 1012 are speech-recognized in parallel, and the utterances 1001 to 1008 are also speech-recognized in parallel. That is, the past utterances and the real-time utterances are speech-recognized in parallel, and the past utterances themselves are also speech-recognized in parallel.
[0069] In the example shown in FIG. 8, at the call time of "00:45", the speech recognition texts of speech recognition using the changed selected dictionary 500 for utterances 1001 to 1002, utterances 1005 to 1006, and utterance 1009 are obtained. In this example, the parallel number is 3, and the past utterances and the real-time utterances are speech-recognized in parallel, and within the past utterances, utterances 1001 to 1002 and utterances 1005 to 1006 are speech-recognized in parallel. Also, at the call time of "00:55", the speech recognition texts of speech recognition using the changed selected dictionary 500 for utterances 1001 to 1012 are obtained.
[0070] In this way, when the currently selected dictionary 500 is changed, in this voice recognition example, the changed currently selected dictionary 500 is used to perform parallel voice recognition on the utterance before the change and the utterance after the change, and also perform parallel voice recognition on the utterance before the change. As a result, for example, it becomes possible to perform real-time voice recognition while simultaneously performing voice recognition on past utterances. Also, for example, it is possible to perform voice recognition on past utterances with priority. Furthermore, since past utterances are recognized in parallel, it is possible to complete the voice recognition of past utterances quickly.
[0071] <Details of the response support screen in step S107 of FIG. 3> Hereinafter, the details of the response support screen in step S107 of FIG. 3 will be described. In step S107 of FIG. 3, either the following response support screen example 1 or response support screen example 2 is displayed on the user terminal 20 as the response support screen.
[0072] · Response support screen example 1 In response support screen example 1, the voice recognition text of the latest real-time utterance is always displayed on the screen. In this case, the voice recognition text of past utterances is visualized in the background.
[0073] For example, FIG. 9 shows the response support screen in the case where voice recognition is performed according to voice recognition example 4 or voice recognition example 5. As shown in FIG. 9, the voice recognition text of the latest real-time utterance (utterance 1009 in the example shown in FIG. 9) is always displayed in the utterance display column 2100 of the response support screen 2000. When a new real-time utterance is made, the utterance display column 2100 is automatically scrolled, and the voice recognition text of that real-time utterance is displayed. On the other hand, the voice recognition text of past utterances is visualized in the background (that is, the non-displayed part of the utterance display column 2100).
[0074] This response support screen example 1 is preferably used, for example, in voice recognition example 1, voice recognition example 4, or voice recognition example 5.
[0075] · Example of Response Support Screen 2 In the example of the response support screen 2, the screen is divided into two parts. On one screen, the speech recognition text of the latest real-time speech is always displayed, and on the other screen, the speech recognition text of past speeches is displayed.
[0076] For example, Fig. 10 shows the response support screen when speech recognition is performed according to Speech Recognition Example 4 or Speech Recognition Example 5. As shown in Fig. 10, the speech recognition text of the latest real-time speech (Speech 1009 in the example shown in Fig. 10) is always displayed in the first speech display column 3100 of the response support screen 3000, and the speech recognition text of past speeches is displayed in the second speech display column 3200. When a new real-time speech is made, the first speech display column 3100 is automatically scrolled, and the speech recognition text of that real-time speech is displayed. On the other hand, the speech recognition text of past speeches (including not only the speech recognition text recognized using the changed selected dictionary 500 but also the speech recognition text that has not yet been recognized using the changed selected dictionary 500) is displayed in the second speech display column 3200.
[0077] This example of the response support screen 2 may be used in any of the speech recognition examples from Speech Recognition Example 1 to Speech Recognition Example 5, for example.
[0078] Regarding the speech recognition text of past speeches displayed in the second speech display column 3200, for example, the latest speech recognition text among the speech recognition texts recognized using the changed selected dictionary 500 may be displayed. Also, for example, when the speech recognition of past speeches using the changed selected dictionary 500 is completed, only the first speech display column 3100 may be displayed (that is, when the speech recognition of past speeches using the changed selected dictionary 500 is completed, the second speech display column 3200 may be made non-displayable).
[0079] <Summary> As described above, in the contact center system 1 according to the present embodiment, when the speech recognition dictionary 500 used for speech recognition of the speech (utterance) of the call between the operator and the customer is changed, the speech recognition is performed again on the utterance before the change using the changed speech recognition dictionary 500. Thereby, even when an appropriate speech recognition dictionary 500 is not selected at the start of the call, it becomes possible to perform speech recognition on the entire call using the appropriate speech recognition dictionary 500. For this reason, it becomes possible to obtain a highly accurate speech recognition result, and as a result, for example, it is possible to contribute to an improvement in response quality, an improvement in the accuracy of various analyses, and the like.
[0080] <Other: Supplementary> · In the above Speech Recognition Example 2 to Speech Recognition Example 5, when the selected dictionary 500 is changed, the speech recognition of past utterances is performed again. Therefore, when the time until the end of the call is short, there is a possibility that the speech recognition may not be completed. Therefore, in such a case, the speech recognition is continued even after the call ends. Thereby, the utterances of the entire call can be speech-recognized using the appropriate speech recognition dictionary 500.
[0081] · When the selected dictionary 500 is changed, which of the above Speech Recognition Example 2 to Speech Recognition Example 5 is used for speech recognition may be fixedly set in advance, or may be set to be changeable by the user (administrator, supervisor, operator, etc.). That is, when the selected dictionary 500 is changed, whether to process the past speech files in parallel for each utterance section or not, and whether to process the past speech files and the real-time speech files in parallel or not may be fixedly set in advance, or may be set to be changeable by the user.
[0082] <Modification Example> Hereinafter, several modification examples of the present embodiment will be described.
[0083] · Modification Example 1 In the above embodiment, when the speech recognition dictionary 500 is changed, the utterance before the change (past utterance) is speech-recognized again by the speech recognition dictionary 500 after the change. However, depending on the relationship between the speech recognition dictionary 500 before the change and the speech recognition dictionary 500 after the change, it may not be necessary to speech-recognize the past utterance again.
[0084] For example, when the speech recognition dictionary 500 before the change is the "speech recognition dictionary 500 specialized for financial operations" and the speech recognition dictionary 500 after the change is the "speech recognition dictionary 500 specialized for insurance operations", it may not be necessary to speech-recognize the past utterance again. This is because it is considered that after inquiries regarding finance were handled within one call, inquiries regarding insurance were then handled, and it is thought that for each inquiry handling, the appropriate speech recognition dictionary 500 was selected by the operator.
[0085] On the other hand, when the speech recognition dictionary 500 before the change is the "speech recognition dictionary 500 for general-purpose use" and the speech recognition dictionary 500 after the change is the "speech recognition dictionary 500 specialized for specific operations", the past utterance is speech-recognized again. This is because it is considered that initially the operator could not select the appropriate speech recognition dictionary 500, and the speech recognition dictionary 500 for general-purpose use was selected as the default dictionary 500, and then the appropriate speech recognition dictionary 500 was selected by the operator.
[0086] In addition to the above, for example, depending on the inquiry content, purpose, the products or technologies targeted by the inquiry, etc., it may not be necessary to speech-recognize the past utterance again. For example, when the product targeted by the inquiry is the same type of insurance, or when the target changes from all financial products to insurance, or in the case of technologies or products in the same field, if the language, vocabulary, etc. used by the speech recognition dictionary 500 before the change are corresponding and inclusive, it may not be necessary to perform speech recognition again using the speech recognition dictionary 500 after the change. Also, when it can be determined from the speech recognition result that the purposes of the inquiries are common or similar, or when it can be understood from the speech recognition dictionary and its attributes, etc. that both the speech recognition dictionary 500 before the change and the speech recognition dictionary 500 after the change can handle it, it may not be necessary to perform speech recognition again using the speech recognition dictionary 500 after the change.
[0087] ·Modification Example 2 In the above embodiment, when the voice recognition dictionary 500 is changed, the past utterances of both the operator and the customer are re-voice recognized by the changed voice recognition dictionary 500. However, only the past utterance of either one (only the past utterance of the customer or only the past utterance of the operator) may be re-voice recognized. For example, when the customer speaks a dialect, only the voice recognition dictionary 500 of the customer may be changed according to the dialect spoken by the customer, and only the utterance of the customer may be re-voice recognized. By having the voice recognition dictionaries 500 independently in this way, the target of re-voice recognition can be limited and the load of re-voice recognition can be reduced.
[0088] ·Modification Example 3 In the above embodiment, the voice recognition dictionary 500 is assumed to be common to the customer and all operators, but it is not limited to this. The voice recognition dictionary 500 that can be selected by the operator may be different according to, for example, the speaking characteristics and business field of the individual operator. That is, each operator may be able to select a voice recognition dictionary 500 suitable for their own speaking characteristics and business field. Also, the voice recognition dictionary 500 of the operator may be selected according to the customer. For example, when the customer speaks a dialect, when the operator speaks with the dialect according to the customer, the voice recognition dictionary 500 of the operator may be changed from a dictionary corresponding only to the standard language to a voice recognition dictionary 500 corresponding to both the dialect spoken by the customer and the standard language in the middle. At this time, only the past utterance of the operator whose voice recognition dictionary 500 has been changed needs to be the target of re-voice recognition. If it can be understood from the attributes of the voice recognition dictionary etc. that the changed voice recognition dictionary 500 can correspond to both the dialect spoken by the customer and the standard language spoken by the operator as described above, it may not be necessary to perform re-voice recognition again.
[0089] The present invention is not limited to the above specifically disclosed embodiments, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.
Explanation of Reference Numerals
[0090] 1 Contact Center System 10 Speech Recognition System 20 User Terminal 30 Telephone 40 PBX 50 NW Switch 60 Customer Terminal 70 Communication Network 101 Voice Recording Unit 102 Dictionary Selection Unit 103 Voice Recognition Unit 104 UI Provision Unit 105 Voice Memory Unit 106 Dictionary Memory Unit 107 Call History Memory Unit 201 UI Control Unit E Contact Center Environment
Claims
1. a selection unit that selects a voice recognition dictionary to be used for voice recognition from among a plurality of voice recognition dictionaries; a voice recognition unit that generates a voice recognition text by converting an utterance included in a voice call with a customer into text by voice recognition using the voice recognition dictionary selected by the selection unit; having When the voice recognition dictionary selected by the selection unit is changed, the information processing system generates a voice recognition text in which some or all of the utterances included in the voice call are converted into text in parallel by the voice recognition using the changed voice recognition dictionary.
2. a selection unit that selects a voice recognition dictionary to be used for voice recognition from among a plurality of voice recognition dictionaries; a voice recognition unit that generates a first voice recognition text by converting an utterance included in a voice call with a customer into text through the voice recognition using the voice recognition dictionary selected by the selection unit; a display unit that displays the first speech recognition text on a screen of a terminal used by a person other than the person who is conducting the voice call with the customer; having The voice recognition unit when the voice recognition dictionary selected by the selection unit is changed, a second voice recognition text is generated by converting an utterance included in the voice call before the change and an utterance after the change into text by the voice recognition using the changed voice recognition dictionary; The display unit is When the voice recognition dictionary selected by the selection unit is changed, the second voice recognition text is displayed on the screen.
3. a selection step of selecting a speech recognition dictionary to be used for speech recognition from among a plurality of speech recognition dictionaries; a voice recognition step of generating a voice-recognized text by converting an utterance included in a voice call with a customer into text by the voice recognition using the voice recognition dictionary selected by the selection step; The computer executes When the voice recognition dictionary selected by the selection procedure is changed, the changed voice recognition dictionary is used to generate a voice recognition text in which some or all of the utterances included in the voice call are converted into text in parallel by the voice recognition.
4. a selection step of selecting a speech recognition dictionary to be used for speech recognition from among a plurality of speech recognition dictionaries; a voice recognition step of generating a first voice-recognized text by converting an utterance included in a voice call with a customer into text by the voice recognition using the voice recognition dictionary selected by the selection step; a display step of displaying the first speech recognition text on a screen of a terminal used by a person other than the person who is conducting the voice call with the customer; The computer executes The speech recognition step includes: When the voice recognition dictionary selected by the selection step is changed, a second voice recognition text is generated by converting an utterance included in the voice call before the change and an utterance after the change into text by the voice recognition using the changed voice recognition dictionary, The display procedure includes: When the voice recognition dictionary selected by the selection step is changed, the second voice recognition text is displayed on the screen.
5. A program for causing a computer to function as the information processing system according to claim 1 or 2.
Citation Information
Patent Citations
Speech recognition device and method used therefor, and record medium where control program therefor is recorded
JP2000200093A
Operator's work support system
JP2006276754A
Voice recognition device, voice recognition method, and program and recording medium therefor
JP2011141349A
Voice recognition device, method thereof, and program
JP2013101204A
Word registering apparatus, and computer program for the same
JP2014048506A