Information processing systems, information processing methods, and programs
The system addresses the issue of inaccurate speech recognition by dynamically selecting and re-selecting dictionaries, ensuring accurate transcription in contact centers.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NTT TECHNOCROSS CORP
- Filing Date
- 2025-03-26
- Publication Date
- 2026-07-24
AI Technical Summary
Conventional speech recognition systems in contact centers often use pre-configured dictionaries that result in insufficient accuracy due to the difficulty in selecting the appropriate dictionary for speech recognition.
An information processing system that includes a selection unit to choose from multiple speech recognition dictionaries and performs speech recognition using the selected dictionary, allowing for re-recognition with a changed dictionary if necessary, to enhance accuracy.
This approach enables highly accurate speech recognition results by ensuring the appropriate dictionary is used throughout the call, improving service quality and analysis accuracy.
Smart Images

Figure 0007894969000001 
Figure 0007894969000002 
Figure 0007894969000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system, an information processing method, and a program.
Background Art
[0002] In speech recognition technology, generally, a speech recognition dictionary in which the notations, pronunciations, arrangements, etc. of words are registered is used. There are various types of such speech recognition dictionaries according to the applications, languages, etc. targeted for speech recognition. For example, there are general-purpose dictionaries, dictionaries in which many technical terms related to specific operations are registered, dictionaries specialized for specific languages, dictionaries specialized for specific dialects, and the like.
[0003] In a contact center (or also called a call center), by a speech recognition system implementing the above speech recognition technology, the speech during a call is converted into text in real time, and the text is presented to an operator (for example, Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, conventionally, even when multiple speech recognition dictionaries were available, it was difficult for operators to select the appropriate dictionary. As a result, speech recognition was performed using a pre-configured speech recognition dictionary for the operator (for example, a general-purpose speech recognition dictionary set as the default), which sometimes resulted in speech recognition results with insufficient accuracy.
[0006] This disclosure is made in view of the above points and aims to provide a technology that can obtain highly accurate speech recognition results. [Means for solving the problem]
[0007] An information processing system according to one aspect of the present disclosure includes a selection unit that selects a speech recognition dictionary to be used for speech recognition from among a plurality of speech recognition dictionaries, and a speech recognition unit that generates speech recognition text by transcribing utterances included in a voice call with a customer into text using the speech recognition dictionary selected by the selection unit, wherein if the speech recognition dictionary selected by the selection unit is changed, the system generates speech recognition text by transcribing some or all of the utterances included in the voice call into text using the changed speech recognition dictionary in parallel. [Effects of the Invention]
[0008] A technology is provided that enables obtaining highly accurate speech recognition results. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of the overall configuration of the contact center system according to this embodiment. [Figure 2] This figure shows an example of the functional configuration of the contact center system according to this embodiment. [Figure 3] This is a sequence diagram showing an example of the customer support processing according to this embodiment. [Figure 4] This is a diagram (part 1) illustrating an example of speech recognition. [Figure 5]This is a diagram (part 2) illustrating an example of speech recognition. [Figure 6] This is a diagram (part 3) illustrating an example of speech recognition. [Figure 7] This is a diagram (part 4) illustrating an example of speech recognition. [Figure 8] This is a diagram (number 5) illustrating an example of speech recognition. [Figure 9] This is a diagram (part 1) illustrating an example of a customer support screen. [Figure 10] This is a diagram (part 2) illustrating an example of a customer support screen. [Modes for carrying out the invention]
[0010] The following describes one embodiment of the present invention. In this embodiment, a contact center system 1 is described that can obtain highly accurate speech recognition results for the audio of a call between an operator and a customer, when it is possible to automatically or manually select a dictionary from among a plurality of speech recognition dictionaries, and is intended for a contact center. However, the contact center is just an example, and the same can be applied to, for example, an office, when it is possible to automatically or manually select a dictionary from among a plurality of speech recognition dictionaries, and it is possible to obtain highly accurate speech recognition results for the audio of a call between a staff member and a customer.
[0011] <Overall configuration of Contact Center System 1> Figure 1 shows an example of the overall configuration of the contact center system 1 according to this embodiment. As shown in Figure 1, the contact center system 1 according to this embodiment includes a voice recognition system 10, a plurality of user terminals 20, a plurality of telephones 30, a PBX (Private Branch eXchange) 40, a network switch 50, and a customer terminal 60. Here, the voice recognition system 10, user terminals 20, telephones 30, PBX 40, and network switch 50 are installed in the contact center environment E, which is the system environment of the contact center. Note that the contact center environment E is not limited to a system environment in the same building, but may be, for example, a system environment in multiple buildings geographically separated.
[0012] The voice recognition system 10 uses packets (voice packets) transmitted from the network switch 50 to record the audio of a call between an operator and a customer as an audio file. The voice recognition system 10 may passively acquire voice packets transmitted from the network switch 50, or it may actively acquire voice data by requesting voice data from the PBX 40 via the network switch 50.
[0013] Furthermore, the speech recognition system 10 performs speech recognition on this audio file and generates text representing the speech recognition result (hereinafter also referred to as speech recognition text). At this time, if the speech recognition dictionary is changed, the speech recognition system 10 uses the changed speech recognition dictionary to perform speech recognition again on audio files that have already been speech-recognized (that is, it performs speech recognition on speech that has already been speech-recognized using the previous speech recognition dictionary as well). This makes it possible, for example, if an inappropriate speech recognition dictionary is changed to an appropriate one, to re-speech speech that has already been speech-recognized using the inappropriate speech recognition dictionary, thereby obtaining a more accurate speech recognition result. The speech recognition system 10 is implemented, for example, by a general-purpose server or a group of servers.
[0014] The user terminal 20 is a terminal such as a PC (Personal Computer) used by a user (operator or supervisor). In the following, the user is mainly assumed to be an operator, but some users may be supervisors. Note that an operator is a person whose main job is to answer calls from customers, etc. On the other hand, a supervisor is a person who monitors the calls of operators and supports the telephone answering operations of those operators when some problem is likely to occur or in response to a request from an operator. Usually, the calls of several to a dozen or so operators are generally monitored by one supervisor.
[0015] On the user terminal 20, a response support screen is displayed where the voice recognition result (voice recognition text) during a call with a customer is visualized in real time. By referring to this response support screen, the operator can also confirm the content of the call with the customer as text.
[0016] The telephone 30 is an IP (Internet Protocol) telephone (fixed IP telephone or mobile IP telephone, etc.) used by the operator.
[0017] The PBX 40 is a telephone exchange (IP-PBX) and is connected to a communication network 70 including a VoIP (Voice over Internet Protocol) network and a PSTN (Public Switched Telephone Network).
[0018] The NW switch 50 relays packets between the telephone 30 and the PBX 40 and captures those packets and sends them to the voice recognition system 10.
[0019] The customer terminal 60 is various terminals such as a smartphone, mobile phone, or fixed phone used by the customer.
[0020] Note that the overall configuration of the contact center system 1 shown in Figure 1 is just one example, and other configurations are possible. For example, in the example shown in Figure 1, the voice recognition system 10 is included in the contact center environment E (i.e., the voice recognition system 10 is on-premise), but all or part of the functions of the voice recognition system 10 may be implemented by cloud services or the like. Similarly, in the example shown in Figure 1, the PBX 40 is an on-premise telephone exchange, but it may also be implemented by cloud services. Furthermore, if the user terminal 20 has telephone functionality, the contact center system 1 does not necessarily need to include a telephone 30.
[0021] <Functional Configuration of Contact Center System 1> Figure 2 shows an example of the functional configuration of the voice recognition system 10 and user terminal 20 included in the contact center system 1 according to this embodiment.
[0022] ≪Voice Recognition System 10≫ As shown in Figure 2, the speech recognition system 10 according to this embodiment includes a speech recording unit 101, a dictionary selection unit 102, a speech recognition unit 103, and a UI provision unit 104. Each of these units is realized, for example, by processing that one or more programs installed in the speech recognition system 10 cause a processor such as a CPU (Central Processing Unit) to execute. The speech recognition system 10 according to this embodiment also includes a speech storage unit 105, a dictionary storage unit 106, and a call history storage unit 107. Each of these units can be realized, for example, by a storage device such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory. However, at least a portion of the storage area of each of these units may be realized by a storage device (such as a database server) that is connected to the speech recognition system 10 in a communicative manner.
[0023] The voice recording unit 101 stores the voice data represented by the packets (voice packets) transmitted from the NW switch 50 as an audio file in the voice storage unit 105.
[0024] The dictionary selection unit 102 selects a speech recognition dictionary 500 to be used for speech recognition from among multiple speech recognition dictionaries 500 stored in the dictionary storage unit 106. A speech recognition dictionary 500 is dictionary information in which, for example, the spelling and pronunciation of words, the order of words, etc., are registered. There are various types of speech recognition dictionaries 500, such as speech recognition dictionaries for general use, speech recognition dictionaries specialized for specific businesses (e.g., finance, insurance, information and communication, etc.), speech recognition dictionaries specialized for specific languages (e.g., Japanese, English, French, etc.), and speech recognition dictionaries specialized for specific dialects (e.g., the dialect of a certain region of Japan, etc.). Hereinafter, the speech recognition dictionary 500 selected by the dictionary selection unit 102 will also be referred to as the "selected dictionary 500".
[0025] The speech recognition unit 103 uses the selected dictionary 500 selected by the dictionary selection unit 102 to perform speech recognition on the audio files stored in the audio storage unit 105 and generate speech recognition text, which is the result of the speech recognition. At this time, the speech recognition unit 103 performs speech recognition on the speech of each speaker (operator, customer) and generates speech recognition text with speaker information and time information. The speech recognition text of a sentence (a single utterance or a single phrase, etc.) is represented in the format, for example, (speaker information, time information, speech recognition text). Such speech recognition text with speaker information and time information can be generated using known speech recognition technology. Speaker information is information indicating the speaker (operator or customer) who uttered the speech corresponding to the speech recognition text, and time information is information indicating the time (date and time) when the speech corresponding to the speech recognition text was uttered. Hereinafter, the speech recognition text will be assumed to have speaker information and time information attached and will be represented in the format, for example, (speaker information, time information, speech recognition text).
[0026] Furthermore, if the selected dictionary 500 is changed, the speech recognition unit 103 will use the changed selected dictionary 500 to perform speech recognition again on speech files that have already been speech-recognized.
[0027] Furthermore, when a call between an operator and a customer ends, for example, the speech recognition unit 103 stores call history information, including the speech recognition text related to that call, in the call history storage unit 107.
[0028] The UI provider unit 104 provides screen information for a customer support screen where the speech recognition text generated by the speech recognition unit 103 is visualized. The screen information is represented by, for example, HTML (Hypertext Markup Language), CSS (Cascading Style Sheets), JavaScript, etc.
[0029] The voice storage unit 105 stores the audio files of the audio represented by the packets (voice packets) transmitted from the NW switch 50.
[0030] The dictionary storage unit 106 stores multiple speech recognition dictionaries 500. Among these multiple speech recognition dictionaries 500, there is a speech recognition dictionary 500 that is selected as the default (standard) (hereinafter referred to as the "default dictionary 500"). The default dictionary 500 is generally a speech recognition dictionary for general use, but for example, if a contact center mainly handles inquiries for a specific business, the default dictionary 500 may be a speech recognition dictionary specialized for that business. Alternatively, for example, if a contact center mainly handles inquiries from customers who speak a specific language, the default dictionary 500 may be a speech recognition dictionary specialized for that language, or if it handles inquiries from customers in a specific region, the default dictionary 500 may be a speech recognition dictionary specialized for the dialect of that region.
[0031] The call history storage unit 107 stores call history information. Call history information includes, for example, at least the call ID and the speech recognition text related to the call with that call ID. The call history information may also include various other pieces of information, such as the date and time of the call, the duration of the call, the ID of the operator who answered the call, the extension number of the operator, the customer's telephone number, and any memo information related to the call.
[0032] ≪User Terminal 20≫ As shown in Figure 2, the user terminal 20 according to this embodiment has a UI control unit 201. The UI control unit 201 is implemented, for example, by a process that one or more programs (such as a web browser) installed on the user terminal 20 cause to be executed by a processor such as a CPU.
[0033] The UI control unit 201 displays various screens, including response support screens, on the display of the user terminal 20. The UI control unit 201 also accepts various user input operations on these screens.
[0034] <Customer support processing> The following describes the process (response support process) of performing speech recognition on the audio of a call between an operator and a customer and displaying the speech recognition results on the response support screen of the user terminal 20, with reference to Figure 3.
[0035] When a call is initiated between the operator and the customer, the voice recording unit 101 of the voice recognition system 10 receives a packet (start packet) indicating that the call has been initiated (step S101).
[0036] Next, the dictionary selection unit 102 of the speech recognition system 10 selects a speech recognition dictionary 500 to be used for speech recognition from among the multiple speech recognition dictionaries 500 stored in the dictionary storage unit 106 (step S102). Here, the dictionary selection unit 102 may, for example, select the default dictionary 500, or it may query the user terminal 20 to determine which speech recognition dictionary 500 to use and then select the speech recognition dictionary 500 specified by the user (operator) in response to this query. Also, when querying the user terminal 20 to determine which speech recognition dictionary 500 to use, the dictionary selection unit 102 may, for example, give the user (operator) a certain grace period of several tens of seconds, and if no speech recognition dictionary 500 is specified within this grace period, it may select the default dictionary 500 (in this case, speech recognition will not be performed until the grace period has elapsed). Generally, it is difficult for the operator to determine which speech recognition dictionary 500 to use at the start of a call. Alternatively, for example, the default dictionary 500 may be considered selected until the operator explicitly selects the speech recognition dictionary 500.
[0037] The following steps S103 to S108 are performed repeatedly during a call between the operator and the customer.
[0038] The voice recording unit 101 of the voice recognition system 10 receives packets (voice packets) transmitted from the NW switch 50 (step S103).
[0039] Next, the voice recording unit 101 of the voice recognition system 10 saves the voice data represented by the packet as an audio file in the voice storage unit 105 (step S104).
[0040] Next, the speech recognition unit 103 of the speech recognition system 10 performs speech recognition on the audio files stored in the audio storage unit 105 using the selected dictionary 500, and generates speech recognition text, which is the result of the speech recognition (step S105). At this time, if the selected dictionary 500 is changed in step S108, which will be described later, the speech recognition unit 103 will use the changed selected dictionary 500 to perform speech recognition again on the audio files that have already been speech-recognized. Details of the speech recognition in this step will be described later.
[0041] Next, the UI provision unit 104 of the speech recognition system 10 transmits the speech recognition text generated in step S105 and screen information for visualizing that speech recognition text to the user terminal 20 (for example, the user terminal 20 used by the operator making the call) (step S106). Here, the UI provision unit 104 may transmit the speech recognition text and screen information to the user terminal 20 each time speech recognition text is generated in step S105, or it may transmit the speech recognition text and screen information to the user terminal 20 in response to a request from the user terminal 20. In addition, the UI provision unit 104 may transmit the speech recognition text and screen information not only to the user terminal 20 used by the operator making the call, but also to, for example, a user terminal 20 used by a supervisor monitoring the operator's call.
[0042] When the UI control unit 201 of the user terminal 20 receives speech recognition text and screen information, it displays the speech recognition text on the support screen based on this screen information (step S107). Details of the support screen in this step will be described later.
[0043] When changing the selected dictionary 500, the dictionary selection unit 102 of the speech recognition system 10 changes the selected dictionary 500 to one of the multiple speech recognition dictionaries 500 (step S108). Here, the dictionary selection unit 102 only needs to change the selected dictionary 500 to the speech recognition dictionary 500 specified by the user (operator), for example. This is because, after a certain amount of conversation has taken place, the operator can determine which speech recognition dictionary 500 to use.
[0044] However, the dictionary selection unit 102 may, by some judgment logic, decide whether or not to change the selected dictionary 500 and which speech recognition dictionary 500 to change it to. For example, the dictionary selection unit 102 may, after identifying what language the call is being made in using known natural language processing, change the selected dictionary 500 to a speech recognition dictionary 500 specialized for the identified language. Similarly, for example, the dictionary selection unit 102 may, after identifying what dialect the customer is speaking using known natural language processing, change the selected dictionary 500 to a speech recognition dictionary 500 specialized for the identified dialect. Alternatively, for example, the dictionary selection unit 102 may, after inferring the content of the business from the frequency of specific words, etc., contained in the speech recognition text to date (for example, speech recognition results using the default dictionary 500, which is a general-purpose speech recognition dictionary 500), use known inference techniques such as machine learning, and then change the selected dictionary 500 to a speech recognition dictionary 500 specialized for that business.
[0045] When a call between an operator and a customer ends, the speech recognition unit 103 of the speech recognition system 10 creates call history information containing the speech recognition text related to that call and stores the call history information in the call history storage unit 107 (step S109). The call history information is used, for example, for various analyses to improve the quality of customer service and for evaluating operators.
[0046] <Details of speech recognition in step S105 of Figure 3> The following describes the details of speech recognition in step S105 of Figure 3. In the following, it is assumed that the default dictionary 500 was selected in step S102 of Figure 3.
[0047] • Example of speech recognition 1: When there is no change in the selected dictionary 500 As shown in Figure 4, it is assumed that at call duration "00:35", the speech recognition text for utterances 1001 to 1008 has been obtained by speech recognition using the default dictionary 500. Utterances 1001, 1003, 1005, and 1007 are the operator's utterances, while utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0048] In this speech recognition example, since the selected dictionary 500 is not changed, the speech recognition text of the operator's utterance 1009 at call duration "00:38" and the speech recognition text of the customer's utterance 1010 at call duration "00:43" are both obtained by speech recognition using the default dictionary 500.
[0049] Similarly, the speech recognition text of operator utterance 1011 at call duration "00:49" and the speech recognition text of customer utterance 1012 at call duration "00:54" are both obtained by speech recognition using the default dictionary 500.
[0050] Thus, if the selected dictionary 500 is not changed, the voice (utterances) during the call will be recognized using that selected dictionary 500.
[0051] • Speech recognition example 2: When the selected dictionary 500 is changed As shown in Figure 5, it is assumed that at call duration "00:35", the speech recognition text for utterances 1001 to 1008 has been obtained by speech recognition using the default dictionary 500. Utterances 1001, 1003, 1005, and 1007 are the operator's utterances, while utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0052] In this case, assume that the selected dictionary 500 was changed after call time "00:35" but before call time "00:38". In this case, in this speech recognition example, utterances 1001 to 1008, which have already been recognized, will be recognized in chronological order using the changed selected dictionary 500. On the other hand, utterances 1009 to 1012, after the change in the selected dictionary 500, will be recognized in chronological order after the speech recognition of utterances 1001 to 1008 has been completed.
[0053] In the example shown in Figure 5, at call duration "00:45", speech recognition text is obtained for utterances 1001 to 1003 using the modified selected dictionary 500. Furthermore, at call duration "00:55", speech recognition text is obtained for utterances 1001 to 1012 using the modified selected dictionary 500.
[0054] In this speech recognition example, if the selected dictionary 500 is changed, the speech recognition will first perform speech recognition on the utterances before the change using the changed selected dictionary 500, in chronological order, and then perform speech recognition on the utterances after the change using the changed selected dictionary 500, also in chronological order. Hereafter, the utterances of the operator and customer before the change in the selected dictionary 500 will be referred to as "past utterances," and the utterances of the operator and customer after the change in the selected dictionary 500 will be referred to as "real-time utterances." Furthermore, the audio file containing the audio of past utterances will be referred to as the "past audio file," and the audio file containing the audio of real-time utterances will be referred to as the "real-time audio file." Note that if the audio of past utterances and the audio of real-time utterances are recorded in the same audio file, then the past audio file and the real-time audio file are the same audio file. However, the audio of past utterances and the audio of real-time utterances may be recorded in different audio files. In this case, the past audio file and the real-time audio file will be different audio files.
[0055] • Speech recognition example 3: When the selected dictionary 500 is changed, and past audio files are processed in parallel on an utterance interval basis. In the second speech recognition example above, past utterances are re-recognized in chronological order using the modified selected dictionary 500. This is because, generally, speech recognition processing requires recognition to be performed sequentially from the beginning of the audio file. On the other hand, by performing a process called voice activity detection (VAD) on the audio file, it becomes possible to perform speech recognition in parallel on an utterance segment basis. Therefore, in this speech recognition example, voice activity detection is performed on past audio files, and then past utterances are recognized in parallel. However, the number of speech recognitions that can be performed in parallel (hereinafter also referred to as the number of parallel processes) depends on the number of speech recognition engines, etc., and is a predetermined number.
[0056] As shown in Figure 6, it is assumed that at call duration "00:35", the speech recognition text for utterances 1001 to 1008 has been obtained by speech recognition using the default dictionary 500. Utterances 1001, 1003, 1005, and 1007 are the operator's utterances, while utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0057] In this case, assume that the selected dictionary 500 was changed after call time "00:35" but before call time "00:38". In this case, in this speech recognition example, utterances 1001 to 1008, which have already been recognized, will be recognized in parallel using the changed selected dictionary 500. On the other hand, utterances 1009 to 1012, after the change in the selected dictionary 500, will be recognized in chronological order after the speech recognition of utterances 1001 to 1008 has been completed.
[0058] In the example shown in Figure 6, at call time "00:45", speech recognition text is obtained for utterances 1001 and utterances 1004-1005 using the modified selected dictionary 500. In this example, the number of parallel processes is 2, and utterances 1001 and utterances 1004-1005 are recognized in parallel. Also, at call time "00:55", speech recognition text is obtained for utterances 1001-1012 using the modified selected dictionary 500.
[0059] In this speech recognition example, when the selected dictionary 500 is changed, the system first uses the changed selected dictionary 500 to perform speech recognition on the utterances before the change in parallel, and then uses the changed selected dictionary 500 to perform speech recognition on the utterances after the change in chronological order. This allows for prioritizing speech recognition of past utterances. For example, it becomes possible to prioritize speech recognition of utterances that are close to real time and utterances that are close to the start of the call. Furthermore, because past utterances are recognized in parallel, speech recognition of past utterances can be completed more quickly.
[0060] In this speech recognition example, speech segment detection was performed using a process called speech segment detection, and speech recognition was performed in parallel on a speech segment basis. However, this is just one example; for example, sentences or phrases could be detected, and speech recognition could be performed in parallel on a sentence-by-sentence or phrase-by-phrase basis.
[0061] • Speech recognition example 4: When the selected dictionary 500 is changed, and when past audio files and real-time audio files are processed in parallel. In the second speech recognition example above, all past utterances are recognized using the modified selection dictionary 500, and then real-time utterances are recognized using the modified selection dictionary 500. In contrast, by recording past utterances and real-time utterances in separate audio files, it is possible to perform parallel speech recognition on both. Therefore, in this speech recognition example, past utterances and real-time utterances are recorded in separate audio files, and speech recognition is performed on both in parallel.
[0062] As shown in Figure 7, it is assumed that at call duration "00:35", the speech recognition text for utterances 1001 to 1008 has been obtained by speech recognition using the default dictionary 500. Utterances 1001, 1003, 1005, and 1007 are the operator's utterances, while utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0063] In this case, assume that the selected dictionary 500 was changed after call time "00:35" but before call time "00:38". In this speech recognition example, using the changed selected dictionary 500, utterances 1001 to 1008, which have already been recognized, are recognized in chronological order, and utterances 1009 to 1012 are also recognized in chronological order. That is, past utterances and real-time utterances are recognized in parallel and in chronological order.
[0064] In the example shown in Figure 7, at call time "00:45", speech recognition text is obtained for utterances 1001-1002 and utterance 1009 using the modified selection dictionary 500. This example shows the case where past utterances 1001-1002 and real-time utterance utterance 1009 are recognized in parallel. Also, at call time "00:55", speech recognition text is obtained for utterances 1001-1012 using the modified selection dictionary 500.
[0065] In this way, when the selected dictionary 500 is changed, this speech recognition example uses the changed selected dictionary 500 to perform speech recognition on the utterances before and after the change in parallel and in chronological order. This makes it possible, for example, to perform speech recognition of real-time utterances while simultaneously recognizing past utterances.
[0066] • Speech recognition example 5: When the selected dictionary 500 is changed, and past audio files are processed in parallel on an utterance interval basis, and past audio files and real-time audio files are processed in parallel. This speech recognition example combines the above speech recognition examples 3 and 4. In other words, in this speech recognition example, past utterances and real-time utterances are recorded in separate audio files, and after detecting the utterance interval in the past audio file, speech recognition is performed in parallel on both the past utterances and the real-time utterances, as well as on the past utterances. However, the number of parallel processes for past utterances depends on the number of speech recognition engines, etc., and is a predetermined number.
[0067] As shown in Figure 8, it is assumed that at call duration "00:35", the speech recognition text for utterances 1001 to 1008 has been obtained by speech recognition using the default dictionary 500. Utterances 1001, 1003, 1005, and 1007 are the operator's utterances, while utterances 1002, 1004, 1006, and 1008 are the customer's utterances.
[0068] In this case, assume that the selected dictionary 500 was changed after call time "00:35" but before call time "00:38". In this case, in this speech recognition example, using the changed selected dictionary 500, utterances 1001 to 1008 and utterances 1009 to 1012, which have already been recognized, are recognized in parallel, and utterances 1001 to 1008 are also recognized in parallel. That is, past utterances and real-time utterances are recognized in parallel, and past utterances themselves are also recognized in parallel.
[0069] In the example shown in Figure 8, at call time "00:45", speech recognition text is obtained for utterances 1001-1002, utterances 1005-1006, and utterance 1009 using the modified selected dictionary 500. In this example, the number of parallel processes is 3, and past utterances and real-time utterances are recognized in parallel, with utterances 1001-1002 and utterances 1005-1006 being recognized in parallel within the past utterances. Furthermore, at call time "00:55", speech recognition text is obtained for utterances 1001-1012 using the modified selected dictionary 500.
[0070] In this way, when the selected dictionary 500 is changed, this speech recognition example uses the changed selected dictionary 500 to perform speech recognition on the utterance before the change and the utterance after the change in parallel, and also performs speech recognition on the utterance before the change in parallel. This makes it possible to perform speech recognition of real-time utterances while simultaneously recognizing past utterances. Furthermore, it is possible to prioritize speech recognition of past utterances. In addition, because past utterances are recognized in parallel, it is possible to complete the speech recognition of past utterances more quickly.
[0071] <Details of the support screen in step S107 of Figure 3> The details of the support screen in step S107 of Figure 3 will be explained below. In step S107 of Figure 3, either support screen example 1 or support screen example 2 below will be displayed on the user terminal 20 as the support screen.
[0072] • Example of a customer support screen (Part 1) In the first example of the customer support screen, the most recent real-time speech recognition text is always displayed on the screen. In this case, past speech recognition text is displayed in the background.
[0073] For example, Figure 9 shows the response support screen when speech recognition is performed using speech recognition example 4 or speech recognition example 5. As shown in Figure 9, the speech display area 2100 of the response support screen 2000 always displays the speech recognition text of the most recent real-time utterance (utterance 1009 in the example shown in Figure 9). When a new real-time utterance is made, the speech display area 2100 automatically scrolls and displays the speech recognition text of that real-time utterance. Meanwhile, the speech recognition text of past utterances is visualized in the background (i.e., the hidden part of the speech display area 2100).
[0074] This example of a support screen (Example 1) is preferably used in, for example, speech recognition example 1, speech recognition example 4, or speech recognition example 5.
[0075] • Example of a customer support screen (part 2) In the second example of the support screen, the screen is divided into two parts: one screen always displays the most recent real-time speech recognition text, and the other screen displays the speech recognition text of past speech.
[0076] For example, Figure 10 shows the response support screen when speech recognition is performed using speech recognition example 4 or speech recognition example 5. As shown in Figure 10, the first utterance display field 3100 of the response support screen 3000 always displays the speech recognition text of the most recent real-time utterance (utterance 1009 in the example shown in Figure 10), and the second utterance display field 3200 displays the speech recognition text of past utterances. When a new real-time utterance is made, the first utterance display field 3100 is automatically scrolled to display the speech recognition text of that real-time utterance. On the other hand, the speech recognition text of past utterances (including not only speech recognition text that has been recognized using the changed selected dictionary 500, but also speech recognition text that has not yet been recognized using the changed selected dictionary 500) is displayed in the second utterance display field 3200.
[0077] This example of a support screen (Example 2) may be used in any of the speech recognition examples (Example 1 to Example 5), for example.
[0078] Furthermore, regarding the speech recognition text of past utterances displayed in the second speech display field 3200, for example, the most recent speech recognition text among those recognized using the modified selected dictionary 500 may be displayed. Also, for example, if speech recognition of past utterances using the modified selected dictionary 500 is completed, only the first speech display field 3100 may be displayed (that is, if speech recognition of past utterances using the modified selected dictionary 500 is completed, the second speech display field 3200 may be hidden).
[0079] <Summary> As described above, in the contact center system 1 according to this embodiment, if the speech recognition dictionary 500 used for speech recognition of the voice (utterances) of a call between an operator and a customer is changed, speech recognition is performed again using the changed speech recognition dictionary 500 even for utterances made before the change. This makes it possible to perform speech recognition on the entire call using the appropriate speech recognition dictionary 500, even if the appropriate speech recognition dictionary 500 was not selected at the start of the call. As a result, it becomes possible to obtain highly accurate speech recognition results, which can contribute to improvements in service quality, accuracy of various analyses, etc.
[0080] <Other: Supplementary Information> In the above speech recognition examples 2 through 5, if the selected dictionary 500 is changed, speech recognition of past utterances will be performed again. Therefore, if the time remaining until the end of the call is short, speech recognition may not be completed. In such cases, speech recognition will continue even after the end of the call. This allows speech recognition of the entire utterance of the call to be performed using the appropriate speech recognition dictionary 500.
[0081] When the selected dictionary 500 is changed, the method used for speech recognition from speech recognition example 2 to speech recognition example 5 may be predetermined and fixed, or it may be set to be changeable by the user (administrator, supervisor, operator, etc.). In other words, when the selected dictionary 500 is changed, whether or not to process past audio files in parallel on an utterance interval basis, and whether or not to process past audio files and real-time audio files in parallel, may be predetermined and fixed, or it may be set to be changeable by the user.
[0082] <Variation> The following describes some variations of this embodiment.
[0083] • Variation 1 In the above embodiment, when the speech recognition dictionary 500 is changed, the utterance before the change (past utterance) is re-recognized using the changed speech recognition dictionary 500. However, depending on the relationship between the speech recognition dictionary 500 before the change and the changed speech recognition dictionary 500, it may not be necessary to re-recognize the past utterance.
[0084] For example, if the original speech recognition dictionary 500 was "Speech recognition dictionary 500 specialized for financial services" and the new speech recognition dictionary 500 is "Speech recognition dictionary 500 specialized for insurance services," it is not necessary to re-recognize past utterances. This is because it is assumed that a financial inquiry was handled before an insurance inquiry was handled within the same call, and that the appropriate speech recognition dictionary 500 was selected by the operator for each type of inquiry.
[0085] On the other hand, if the speech recognition dictionary 500 before the change is a "general-purpose speech recognition dictionary 500" and the speech recognition dictionary 500 after the change is a "speech recognition dictionary 500 specialized for a specific task," then past utterances will be re-recognized. This is because it is assumed that initially the operator was unable to select the appropriate speech recognition dictionary 500, so the general-purpose speech recognition dictionary 500 was selected as the default dictionary 500, and then the appropriate speech recognition dictionary 500 was selected by the operator afterward.
[0086] In addition to the above, depending on the content and nature of the inquiry, and the products or technologies to which the inquiry pertains, it may not be necessary to perform speech recognition again on past utterances. For example, if the product to which the inquiry pertains is the same type of insurance, or if the subject shifts from financial products in general to insurance, or if it concerns technologies or products in the same field, and the language and vocabulary used by the previous speech recognition dictionary 500 are compatible or encompassed, it may not be necessary to perform speech recognition again using the revised speech recognition dictionary 500. Furthermore, if it can be determined from the speech recognition results that the purpose of the inquiry is common or similar, and it is clear from the speech recognition dictionary and its attributes that both the previous and revised speech recognition dictionary 500 can handle it, it may not be necessary to perform speech recognition again using the revised speech recognition dictionary 500.
[0087] • Variation 2 In the above embodiment, when the speech recognition dictionary 500 was changed, the past utterances of both the operator and the customer were re-recognized using the changed speech recognition dictionary 500. However, it is also possible to re-recognize only the past utterances of either the customer or only the operator. For example, if the customer speaks in a dialect, only the customer's speech recognition dictionary 500 may be changed according to the dialect spoken by the customer, and only the customer's utterances may be re-recognized. By having separate speech recognition dictionaries 500 in this way, the target of re-recognition can be limited, and the load of re-recognition can be reduced.
[0088] • Modification example 3 In the above embodiment, the speech recognition dictionary 500 is assumed to be common to both customers and all operators, but it is not limited to this. The speech recognition dictionary 500 that an operator can select may differ depending on, for example, the individual operator's speech characteristics or field of work. That is, each operator may be able to select a speech recognition dictionary 500 that is suitable for their own speech characteristics or field of work. In addition, the operator's speech recognition dictionary 500 may be selected according to the customer. For example, when a customer speaks in dialect, and the operator speaks with the customer incorporating the dialect, the operator's speech recognition dictionary 500 may be changed midway through from a dictionary that only supports standard Japanese to a speech recognition dictionary 500 that supports both the customer's dialect and standard Japanese. In this case, only the past utterances of the operator whose speech recognition dictionary 500 has been changed should be targeted for re-speech recognition, and if it can be seen from the attributes of the speech recognition dictionary, etc., that the changed speech recognition dictionary 500 is capable of supporting both the customer's dialect and the operator's standard Japanese as described above, then re-speech recognition is not necessary.
[0089] The present invention is not limited to the embodiments specifically disclosed above, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims. [Explanation of Symbols]
[0090] 1. Contact Center System 10. Voice Recognition System 20 User Terminals 30 telephone 40 PBX 50 NW Switches 60 Customer terminals 70 Communication Networks 101 Audio Recording Department 102 Dictionary Selection Section 103 Voice Recognition Unit 104 UI provision department 105 Voice memory unit 106 Dictionary Memory Section 107 Call history storage unit 201 UI Control Unit E Contact Center Environment
Claims
1. A selection unit that selects a speech recognition dictionary to be used for speech recognition from among multiple speech recognition dictionaries, A speech recognition unit generates speech recognition text by transcribing utterances contained in a voice call with a customer into text using the speech recognition unit, using the speech recognition dictionary selected by the selection unit. It has, The aforementioned speech recognition unit, An information processing system that, when the speech recognition dictionary selected by the selection unit is changed, uses the changed speech recognition dictionary to generate speech recognition text, which is obtained by transcribing past utterances in the voice call and real-time utterances in the voice call in parallel using speech recognition.
2. The voice recognition unit is The information processing system according to claim 1, which uses the modified speech recognition dictionary to perform speech recognition on the past utterances in parallel and generate the speech recognition text based on a priority corresponding to at least one or both of the difference between the utterance time of the past utterance and the start time and the current time of the voice call.
3. A selection procedure for choosing a speech recognition dictionary to be used for speech recognition from among multiple speech recognition dictionaries, A speech recognition procedure that generates speech recognition text by transcribing utterances contained in a voice call with a customer into text using the speech recognition dictionary selected by the selection procedure, The computer executes this, The aforementioned speech recognition procedure is: An information processing method that, if the speech recognition dictionary selected by the above selection procedure is changed, generates speech recognition text by using the changed speech recognition dictionary to convert past utterances in the voice call and real-time utterances in the voice call into text in parallel using the speech recognition method.
4. The speech recognition procedure is: The information processing method according to claim 3, wherein the modified speech recognition dictionary is used to perform speech recognition on the past utterances in parallel and generate the speech recognition text based on a priority corresponding to at least one or both of the difference between the utterance time of the past utterance and the start time of the voice call and the current time.
5. A program that causes a computer to function as the information processing system described in claim 1 or 2.