Speech recognition system, speech recognition method, and program
The voice recognition system optimizes resource usage by only performing real-time voice recognition when the results are being actively referred to, addressing the inefficiency in conventional systems where recognition is performed for all calls regardless of usage.
Patent Information
- Application Number
- JP2023576298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-01-25
AI Technical Summary
Conventional voice recognition systems perform real-time voice recognition for all calls, even when the recognition results are not being referred to, leading to wastage of resources such as CPU resources.
A voice recognition system that includes a voice recognition control unit to determine whether to perform real-time voice recognition, a voice recognition unit for performing the recognition, and a UI providing unit for displaying the recognition results. The system only performs real-time voice recognition when the recognition results are being referred to on a terminal.
This approach improves the efficiency of resource usage for voice recognition by only allocating resources when the recognition results are actively being used, thereby reducing unnecessary processing and conserving CPU resources.
Smart Images

Figure 0007693847000001 
Figure 0007693847000002 
Figure 0007693847000003
Abstract
Description
Technical Field
[0001] The present invention relates to a voice recognition system, a voice recognition method, and a program.
Background Art
[0002] A voice recognition system that records voice during a call and converts it into text in real time for a contact center (also called a call center) has been conventionally known (for example, Non-Patent Document 1). In such a voice recognition system, it is common to perform voice recording and voice recognition for all calls in the contact center.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, conventionally, voice recognition has been performed in real time even for calls that do not necessarily require real-time voice recognition. For example, voice recognition has been performed in real time even when the voice recognition result is not being referred to by anyone, such as when the operator has not launched a UI (user interface) for confirming the voice recognition result. For this reason, resources (especially CPU (Central Processing Unit) resources, etc.) have been wasted.
[0005] One embodiment of the present invention has been made in view of the above points, and an object thereof is to improve the efficiency of resources used for voice recognition.
Means for Solving the Problem
[0006] To achieve the above object, a voice recognition system according to an embodiment includes a voice recognition control unit configured to determine whether to perform voice recognition on voice data acquired from a voice call in real time, a voice recognition unit configured to perform the voice recognition on the voice data determined to perform voice recognition in real time and create a text representing the result of the voice recognition, and a UI providing unit configured to display a screen on which the text can be referred to in real time on a terminal connected via a communication network. The voice recognition control unit is configured to determine to perform voice recognition on the voice data that is the source of the text that can be referred to on the screen in real time when the screen is displayed on the terminal.
Advantages of the Invention
[0007] It is possible to improve the efficiency of the resources used for voice recognition.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0009] Hereinafter, an embodiment of the present invention will be described. In this embodiment, a contact center system 1 that can improve the use resources (particularly, CPU resources, etc.) for speech recognition of the voice recorded from the operator's call will be described for the contact center. However, the contact center is an example, and it can be similarly applied to cases where the use resources for speech recognition of the voice recorded from the call of a person in charge working in, for example, an office are improved. More generally, it can be similarly applied to cases where the use resources for speech recognition of the voice recorded from a certain call are improved.
[0010] <Overall Configuration of Contact Center System 1> An example of the overall configuration of the contact center system 1 according to this embodiment is shown in FIG. 1. As shown in FIG. 1, the contact center system 1 according to this embodiment includes a speech recognition system 10, a plurality of terminals 20, a plurality of telephones 30, a PBX (Private Branch eXchange) 40, an NW switch 50, and a customer terminal 60. Here, the speech recognition system 10, the terminal 20, the telephone 30, the PBX 40, and the NW switch 50 are installed in a contact center environment E, which is a system environment of the contact center. Note that the contact center environment E is not limited to a system environment within the same building, and may be, for example, a system environment within a plurality of geographically separated buildings.
[0011] The voice recognition system 10 records the voice of the call between the operator and the customer using the packets (voice packets) transmitted from the NW switch 50. Further, the voice recognition system 10 performs voice recognition on the recorded voice and converts it into text (hereinafter also referred to as "call text"). At this time, when the call text is referred to in real time by the operator or the supervisor, the voice recognition system 10 performs voice recognition on the voice of the call between the operator and the customer in real time, and if not, does not perform the voice recognition in real time. Note that a supervisor is, for example, a person who monitors the operator's calls and supports the operator's telephone response work in case some problem is likely to occur or in response to a request from the operator. Usually, it is common for the calls of several to a dozen or so operators to be monitored by one supervisor.
[0012] Hereinafter, the screen for the operator or the supervisor to refer to the call text in real time will be referred to as the "real-time call text screen". The real-time call text screen displays the call text, which is the result of the voice recognition performed in real time, in real time.
[0013] The terminal 20 is various terminals such as a PC (personal computer) used by the operator or the supervisor. Hereinafter, the terminal 20 used by the operator will be referred to as the "operator terminal 21", and the terminal 20 used by the supervisor will be referred to as the "supervisor terminal 22".
[0014] The telephone 30 is an IP (Internet Protocol) telephone (such as a fixed IP telephone or a mobile IP telephone) used by the operator. Generally, one operator terminal 21 and one telephone 30 are installed at the operator's seat.
[0015] PBX40 is a private branch exchange (IP-PBX) and is connected to a communication network 70 including a VoIP (Voice over Internet Protocol) network and a PSTN (Public Switched Telephone Network).
[0016] The NW switch 50 relays packets between the telephone 30 and the PBX40 and captures the packets and transmits them to the voice recognition system 10.
[0017] The customer terminal 60 is various terminals such as a smartphone, a mobile phone, and a landline phone used by the customer.
[0018] Note that the overall configuration of the contact center system 1 shown in FIG. 1 is an example, and other configurations may be used. For example, in the example shown in FIG. 1, the PBX40 is an on-premises private branch exchange, but it may also be a private branch exchange realized by a cloud service. Also, for example, the voice recognition system 10 may be realized by a single server and may be called a voice recognition device. Further, when the operator terminal 21 also functions as an IP phone, the operator terminal 21 and the telephone 30 may be integrally configured.
[0019] <Real-time call text screen> An example of the real-time call text screen is shown in FIG. 2. The real-time call text screen 1000 shown in FIG. 2 includes a real-time call text display column 1100, and each time voice recognition is performed in real time by the voice recognition system 10, the call text obtained by the voice recognition is displayed in real time in the real-time call text display column 1100 (that is, the call text obtained by the voice recognition is immediately displayed in the real-time call text display column 1100).
[0020] For example, in the example shown in FIG. 2, call texts 1101 to 1106 are displayed in the real-time call text display column 1100.
[0021] As a result, an operator or supervisor can confirm the conversation between the current operator and the customer in real time by referring to the real-time call text screen.
[0022] <Functional configurations of the voice recognition system 10 and the terminal 20> An example of the functional configurations of the voice recognition system 10 and the terminal 20 according to this embodiment is shown in FIG. 3.
[0023] ≪Voice recognition system 10≫ As shown in FIG. 3, the voice recognition system 10 according to this embodiment includes a recording unit 101, a voice recognition control unit 102, a voice recognition unit 103, a search unit 104, and a UI providing unit 105. Each of these units is realized, for example, by a process in which one or more programs installed in the voice recognition system 10 cause a processor such as a CPU to execute. Further, the voice recognition system 10 according to this embodiment includes a voice data storage unit 106, a call data storage unit 107, a call list storage unit 108, and a display list storage unit 109. Each of these storage units is realized, for example, by an auxiliary storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Note that at least a part of these storage units may be realized by a storage device or the like connected to the voice recognition system 10 via a communication network.
[0024] The recording unit 101 records the voice data included in the voice packet transmitted from the NW switch 50. That is, the recording unit 101 stores the voice data included in the voice packet in the voice data storage unit 106 in association with the call ID. Note that the call ID is information that uniquely identifies a call between an operator and a customer.
[0025] In addition, when a call between an operator and a customer is started, the recording unit 101 adds a pair of the user ID of the operator conducting this call and the call ID of this call to the call list. Further, when the call ends, the recording unit 101 deletes the pair of the user ID of the operator who made this call and the call ID of this call from the call list. Here, the call list is a list in which pairs of the user IDs of the operators currently on call and the call IDs of their calls are stored. Note that the user ID is information that uniquely identifies an operator (and a supervisor).
[0026] The voice recognition control unit 102 controls whether to perform real-time voice recognition (i.e., immediate voice recognition) on the call between the operator and the customer. That is, for calls in which call text is displayed in real time on the real-time call text screen, the voice recognition control unit 102 performs real-time voice recognition on the voice of that call, while for calls that are not like this, it controls to perform voice recognition on the voice of that call in the background at some timing instead of in real time. Also, when the voice recognition control unit 102 performs real-time voice recognition on the voice of a new call, if there is a shortage of CPU resources or the like, it also performs control to suspend part or all of the voice recognition in the background and prioritize real-time voice recognition.
[0027] The voice recognition unit 103 performs voice recognition on the voice data according to the control by the voice recognition control unit 102 to create call text. Also, the voice recognition unit 103 creates call data including at least the call ID and the call text, and stores it in the call data storage unit 107.
[0028] The search unit 104 searches the call data stored in the call data storage unit 107 based on the search conditions received from the UI providing unit 105.
[0029] The UI providing unit 105 provides the terminal 20 with information (hereinafter also referred to as UI information) for displaying the UI (user interface) of various screens (for example, a real-time call text screen, a search screen for the user to specify the above search conditions, etc.) on the terminal 20. Note that the UI information only needs to be information necessary for displaying the screen, and examples include screen definition information in which the screen is defined by HTML (Hypertext Markup Language) or the like.
[0030] Also, when the UI providing unit 105 receives a display request for the real-time call text screen from the terminal 20, it adds the set of user IDs included in the display request to the display list. Further, when the display of the real-time call text screen ends, the UI providing unit 105 deletes the set of user IDs included in the end notification from the display list. Here, the display list is a list in which a pair of the user ID of the operator making a call in which call text is displayed in real time on the real-time call text screen and the user ID of the user (operator or supervisor) of the terminal 20 on which the real-time call text screen is displayed is stored.
[0031] The voice data storage unit 106 stores the voice data recorded by the recording unit 101.
[0032] The call data storage unit 107 stores call data. The call data includes at least a call ID and call text, but may also include various other information such as the calling party phone number and the called party phone number related to the call with the call ID, the user ID of the operator making the call, the call start time and the call end time of the call, and the like.
[0033] The call list storage unit 108 stores a call list in which a pair of the user ID of the operator currently on the call and the call ID of the call is stored.
[0034] The display list storage unit 109 stores a display list in which pairs of the user ID of the operator who is conducting a call in which call text is displayed in real time on the real-time call text screen and the user ID of the user of the terminal 20 on which the real-time call text screen is displayed are stored.
[0035] ≪Terminal 20≫ As shown in FIG. 3, the terminal 20 according to the present embodiment has a UI unit 201. The UI unit 201 is realized, for example, by processing that causes one or more programs installed in the terminal 20 to be executed by a processor such as a CPU.
[0036] The UI unit 201 displays various screens (for example, a real-time call text screen, a search screen, etc.) on a display or the like based on the UI information provided from the UI providing unit 105 of the voice recognition system 10. Further, the UI unit 201 accepts various operations on the screen displayed on the display or the like.
[0037] <Processing of the contact center system 1> Hereinafter, various processes executed by the contact center system 1 according to the present embodiment will be described.
[0038] ≪Display start process of the real-time call text screen≫ The display start process of the real-time call text screen according to the present embodiment will be described with reference to FIG. 4. Hereinafter, the case where a certain user (operator or supervisor) displays a real-time call text screen on the display of his or her own terminal 20 will be described.
[0039] In addition, when the real-time call text screen is not displayed, the real-time call text screen can be displayed at any timing (that is, this process can be started at any timing). Therefore, for example, when it is desired to display a real-time call text screen on the terminal 20 in which the call text of a certain operator's call is displayed in real time, the user (the operator himself or a supervisor who monitors the operator's call) can display the real-time call text screen before the start of the call, or can also display the real-time call text screen during the call.
[0040] First, in response to an operation for displaying the real-time call text screen, the UI unit 201 of the terminal 20 transmits a display request for the real-time call text screen to the voice recognition system 10 (step S101). Here, the display request includes the user ID of the operator (hereinafter also referred to as the display target user ID) for which call text is to be displayed in real time on the real-time call text screen, and the user ID of the user who uses the terminal 20 that transmitted the display request (hereinafter also referred to as the display user ID). When the terminal 20 is the operator terminal 21, the display target user ID and the display user ID are the user ID of the operator who uses the operator terminal 21. On the other hand, when the terminal 20 is the supervisor terminal 22, the display target user ID is the user ID of a certain operator monitored by the supervisor terminal 22, and the display user ID is the user ID of the supervisor who uses the supervisor terminal 22.
[0041] When the UI providing unit 105 of the voice recognition system 10 receives a display request for the real-time call text screen, it adds the display target user ID and the display user ID included in the display request to the display list (step S102).
[0042] Next, the UI providing unit 105 of the voice recognition system 10 transmits the UI information of the real-time call text screen to the terminal 20 (step S103).
[0043] When the UI unit 201 of the terminal 20 receives the UI information of the real-time call text screen, it displays the real-time call text screen on the display based on the UI (step S104).
[0044] <<Display termination process of real-time call text screen>> The display termination process of the real-time call text screen according to this embodiment will be described with reference to FIG. 5. Hereinafter, the case where a certain user (operator or supervisor) terminates the display of the real-time call text screen displayed on the display of his / her own terminal 20 will be described.
[0045] When the real-time call text screen is being displayed, the real-time call text screen can be terminated at any timing (that is, this process can be started at any timing). Therefore, for example, when a real-time call text screen in which the call text of a certain operator's call is displayed in real time is displayed on the terminal 20, the user (the operator himself / herself or the supervisor who monitors the operator's call) can terminate the display of the real-time call text screen during the call, or can also terminate the display of the real-time call text screen after the call ends.
[0046] First, the UI unit 201 of the terminal 20 terminates the display of the real-time call text screen in response to an operation for terminating the display of the real-time call text screen (step S201).
[0047] Next, the UI unit 201 of the terminal 20 transmits a display end notification to the voice recognition system 10 (step S202). Here, the display end notification includes the display target user ID and the display user ID. When the terminal 20 is the operator terminal 21, the display target user ID and the display user ID are the user IDs of the operator who uses the operator terminal 21. On the other hand, when the terminal 20 is the supervisor terminal 22, the display target user ID is the user ID of a certain operator monitored by the supervisor terminal 22, and the display user ID is the user ID of the supervisor who uses the supervisor terminal 22.
[0048] When the UI providing unit 105 of the voice recognition system 10 receives the display end notification, it deletes the display target user ID and the display user ID included in the display end notification from the display list (step S203).
[0049] ≪Processing from Call Start to Call End≫ The processing from call start to call end according to this embodiment will be described with reference to FIG. 6. Hereinafter, the processing from call start to call end of a certain operator will be described.
[0050] First, the recording unit 101 of the voice recognition system 10 receives a call start packet from the NW switch 50 (step S301).
[0051] Next, the recording unit 101 of the voice recognition system 10 adds the user ID included in the call start packet (hereinafter also referred to as the in-call user ID) and the call ID of the call for which the call has started to the call list (step S302). The call ID is arbitrarily generated by the recording unit 101. For example, since one operator can make only one call at a time, the call ID may be generated by adding the call start date and time, etc. to the in-call user ID.
[0052] The following steps S303 to S315 are repeatedly executed during a call (that is, until the recording unit 101 receives a call end packet). Hereinafter, steps S303 to S315 in one repetition will be described.
[0053] The recording unit 101 of the voice recognition system 10 receives a voice packet from the NW switch 50 (step S303). Here, the voice packet includes voice data and a user ID (in-call user ID). At this time, the recording unit 101 identifies the call ID corresponding to the in-call user ID from the call list, and then associates the voice data with the identified call ID and stores it in the voice data storage unit 106.
[0054] The recording unit 101 of the voice recognition system 10 transmits the in-call user ID included in the voice packet received from the NW switch 50 to the voice recognition control unit 102 (step S304).
[0055] When the voice recognition control unit 102 of the voice recognition system 10 receives the in-call user ID, it determines whether it is necessary to perform real-time voice recognition on the voice data of the call ID corresponding to the in-call user ID in the call list (step S305). Specifically, the voice recognition control unit 102 determines whether the in-call user ID is included as the display target user ID in the display list. And when the in-call user ID is included as the display target user ID in the display list, the voice recognition control unit 102 determines that it is necessary to perform real-time voice recognition on the voice data of the call ID corresponding to the in-call user ID in the call list, and otherwise determines that it is not necessary to perform real-time voice recognition on the voice data. Note that when the in-call user ID is included as the display target user ID in the display list, it means that the call text of the call made by the operator of the in-call user ID is referred to in real time on the real-time call text screen.
[0056] If it is determined in step S305 above that real-time speech recognition is necessary, the following steps S306 to S315 are executed.
[0057] The speech recognition control unit 102 of the speech recognition system 10 determines whether there is free space in the resources available for speech recognition (especially CPU resources, etc.) (step S306). Here, the resources available for speech recognition are often represented by an index value called multiplicity, which represents the number of speech data that can be recognized simultaneously. For example, if the multiplicity is N, it means that N speech data can be recognized simultaneously. Therefore, the speech recognition control unit 102 can determine, for example, that there is free space in the resources if the number of speech data currently being recognized simultaneously is n and n < N, and determine that there is no free space in the resources otherwise.
[0058] If it is determined in step S306 above that there is no free space in the resources, the following steps S307 to S309 are executed.
[0059] The speech recognition control unit 102 of the speech recognition system 10 determines the speech data for which speech recognition is to be aborted from the speech data stored in the speech data storage unit 106 according to the following procedures 1 to 3 (step S307).
[0060] Procedure 1: The speech recognition control unit 102 identifies the speech data currently being recognized among the speech data stored in the speech data storage unit 106.
[0061] Procedure 2: Next, the speech recognition control unit 102 identifies the speech data other than the speech data during real-time speech recognition among the speech data identified in Procedure 1. Here, the speech data during real-time speech recognition can be identified by identifying the in-call user ID included as the display target user ID in the display list and then identifying the call ID corresponding to these in-call user IDs from the call list, and then identifying the speech data associated with these call IDs.
[0062] Step 3: Then, the voice recognition control unit 102 determines one or more pieces of voice data from among the voice data identified in Step 2 as voice data for aborting voice recognition. Note that the voice data for aborting voice recognition may be one piece or a plurality of pieces. Also, it may be determined randomly from among the voice data identified in Step 2, or it may be determined according to some criterion. Examples of such a criterion include preferentially determining voice data of a call with a shorter (or longer) elapsed time since the start of voice recognition, preferentially determining voice data of a call of a specific operator (or an operator belonging to a specific group), determining by a round-robin method, and the like.
[0063] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID associated with the voice data determined to abort voice recognition in step S307 above to the voice recognition unit 103 (step S308).
[0064] The voice recognition unit 103 of the voice recognition system 10 aborts the voice recognition of the voice data associated with the call ID received from the voice recognition control unit 102 (step S309). As a result, there is now free space in the resources available for voice recognition.
[0065] When it is determined in step S306 above that there is free space in the resources, or following step S309 above, the voice recognition control unit 102 of the voice recognition system 10 identifies the call ID corresponding to the in-call user ID transmitted from the recording unit 101 in step S304 above from the call list, and transmits the identified call ID and the in-call user ID to the voice recognition unit 103 (step S310).
[0066] The voice recognition unit 103 of the voice recognition system 10 performs voice recognition on the voice data associated with the call ID received from the voice recognition control unit 102 (step S311). As a result, call text, which is the result of performing voice recognition on the voice data, is created.
[0067] Note that, for example, during a call, a real-time call text screen for referring to the call text of the call may be displayed on a certain terminal 20. In this case, there may be no call text until the real-time call text screen is displayed. For a specific example, regarding a call started at time t s if a real-time call text screen for referring to the call text of the call is displayed at a certain time t (> t s ), there may be no call text from time t s to t. In this case, in step S311 described above, the voice recognition unit 103 may perform voice recognition not only on the voice data after time t but also on past voice data (that is, for example, the voice data from time t s to t) at the same time.
[0068] The voice recognition unit 103 of the voice recognition system 10 transmits the call text created in step S311 and the in-call user ID received from the voice recognition control unit 102 in step S310 to the UI providing unit 105 (step S312).
[0069] Also, the voice recognition unit 103 of the voice recognition system 10 associates the call text created in step S311 with the call ID and stores it as call data in the call data storage unit 107 (step S313). At this time, various information such as the in-call user ID may be included in the call data.
[0070] When the UI providing unit 312 of the voice recognition system 10 receives the call text and the in-call user ID, it specifies the display user ID corresponding to the display target user ID that matches the in-call user ID from the display list, and transmits the call text to the terminal 20 of the specified display user ID (step S314).
[0071] When the UI unit 201 of the terminal 20 receives the call text from the voice recognition system 10, it displays the call text on the real-time call text screen (step S315). As a result, the call text is displayed in real time on the real-time call text screen.
[0072] When a call end packet is transmitted from the NW switch 50, the recording unit 101 of the voice recognition system 10 receives the call end packet from the NW switch 50 (step S316).
[0073] Then, the recording unit 101 of the voice recognition system 10 deletes the in-call user ID that matches the user ID included in the call end packet and the corresponding call ID from the call list (step S317).
[0074] ≪Background Voice Recognition Processing≫ The background voice recognition processing according to this embodiment will be described with reference to FIG. 7. This background voice recognition processing is a process for performing voice recognition on voice data other than the voice data that is the target of real-time voice recognition, and is repeatedly executed at predetermined time intervals (for example, every 10 minutes, etc.) in the background of the above-mentioned "real-time call text screen display start processing", "real-time call text screen display end processing", and "processing from call start to call end". However, the time interval for repeating the background voice recognition processing may vary according to, for example, the time zone. For example, the time interval for repetition may be lengthened during the daytime when the call volume is high to perform more real-time voice recognition, and shortened during the nighttime when the call volume is low to perform more voice recognition in the background. Or, the background voice recognition processing may not be executed during the daytime when the call volume is high to perform more real-time voice recognition.
[0075] First, the voice recognition control unit 102 of the voice recognition system 10 determines whether there is free space in the resources (especially CPU resources, etc.) available for voice recognition, in the same manner as step S306 in FIG. 6 (step S401).
[0076] If it is determined in step S401 above that there is no free space in the resources, the following steps S402 to S404 are executed.
[0077] The voice recognition control unit 102 of the voice recognition system 10 determines the voice data to be subjected to voice recognition from among the voice data stored in the voice data storage unit 106 according to the following procedures 11 to 12 (step S402).
[0078] Procedure 11: The voice recognition control unit 102 identifies the voice data that is not currently being subjected to voice recognition among the voice data stored in the voice data storage unit 106.
[0079] Procedure 12: Then, the voice recognition control unit 102 determines one or more pieces of voice data from among the voice data identified in procedure 11 as the voice data to be subjected to voice recognition. Note that the voice data to be subjected to voice recognition may be one piece, or may be multiple pieces depending on the free space status of the resources available for voice recognition. Also, it may be determined randomly from among the voice data identified in procedure 11, or may be determined according to some criteria. Examples of such criteria include preferentially determining the voice data of a call for which the elapsed time since the start of voice recognition is long (or short), preferentially determining the voice data of a call of a certain specific operator (or an operator belonging to a certain specific group), determining by a round-robin method, and the like.
[0080] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID associated with the voice data determined to be subjected to voice recognition in step S402 above to the voice recognition unit 103 (step S403).
[0081] The voice recognition unit 103 of the voice recognition system 10 performs voice recognition on the voice data associated with the call ID received from the voice recognition control unit 102 (step S404). As a result, a call text, which is the result of performing voice recognition on the voice data, is created.
[0082] The voice recognition unit 103 of the voice recognition system 10 associates the call text created in step S404 above with the call ID and stores it as call data in the call data storage unit 107 (step S405). At this time, various information such as the user ID of the operator who made the call with this call ID may be included in the call data.
[0083] ≪Search Process≫ The search process according to this embodiment will be described with reference to FIG. 8. Hereinafter, the case where a certain user (operator or supervisor) searches for call data using his or her own terminal 20 will be described.
[0084] Note that the search for call data can be performed at any timing (that is, this process can be started at any timing).
[0085] The UI unit 201 of the terminal 20 transmits a search request including search conditions specified by the user to the voice recognition system 10 (step S501). Here, as the search conditions, any conditions for searching for call data can be specified. For example, the user ID, call start date and time, call end date and time, call duration, etc. can be specified. The user can specify the search conditions on a search screen for specifying the search conditions, for example.
[0086] When the UI providing unit 105 of the voice recognition system 10 receives the search request from the terminal 20, it transmits the search request to the search unit 104 (step S502).
[0087] When the search unit 104 of the voice recognition system 10 receives a search request from the UI providing unit 105, it searches for the call data stored in the call data storage unit 107 based on the search conditions included in the search request (step S503).
[0088] The search unit 104 of the voice recognition system 10 transmits the search result in step S503 above to the UI providing unit 105 (step S504). Note that the search result includes, for example, the call data searched in step S503 above.
[0089] When the UI providing unit 105 of the voice recognition system 10 receives the search result from the search unit 104, it transmits the search result to the terminal 20 (step S505).
[0090] When the UI unit 201 of the terminal 20 receives the search result from the voice recognition system 10, it displays a search result list that is a list of the call data included in the search result (step S506). The user can select the call data for which they desire detailed display from this search list. Note that this search result list may be displayed on the search screen or on a screen different from the search screen.
[0091] The UI unit 201 of the terminal 20 accepts the selection of the call data to be displayed in detail from the search result list (step S507).
[0092] Here, when voice recognition for the voice data of the call represented by the call data selected by the user is completed, the call data includes the call text of the entire call. On the other hand, when voice recognition for the voice data of the call represented by the call data selected by the user is not completed, the call data either does not include the call text or includes only the call text of a part of the call. Therefore, when voice recognition for the voice data of the call represented by the call data selected by the user is not completed, the following steps S508 to S519 are executed, and when not, the following step S520 is executed. Note that whether the call text is only a part of the call can be determined from, for example, the call time and the like.
[0093] The UI unit 201 of the terminal 20 transmits a voice recognition request to the voice recognition system 10 (step S508). Here, the voice recognition request includes the call ID of the call data selected by the user.
[0094] When the UI providing unit 105 of the voice recognition system 10 receives a voice recognition request from the terminal 20, it transmits the voice recognition request to the voice recognition control unit 102. (Step S509).
[0095] The voice recognition control unit 102 of the voice recognition system 10 determines whether there is free space in the resources available for voice recognition (especially CPU resources, etc.) in the same manner as step S306 in FIG. 6 (step S510).
[0096] If it is determined in step S510 above that there is free space in the resources, the following steps S511 to S516 are executed.
[0097] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID included in the voice recognition request received from the UI providing unit 105 to the voice recognition unit 103 (step S511).
[0098] The speech recognition unit 103 of the speech recognition system 10 performs speech recognition on the speech data associated with the call ID received from the speech recognition control unit 102 (step S512). As a result, call text, which is the result of performing speech recognition on the speech data, is created.
[0099] The speech recognition unit 103 of the speech recognition system 10 transmits the call text created in step S512 above to the UI providing unit 105 (step S513).
[0100] Also, the speech recognition unit 103 of the speech recognition system 10 stores the call text created in step S512 above in the call data storage unit 107 as call data in association with the call ID (step S514). At this time, various information such as the in-call user ID may be included in the call data.
[0101] When the UI providing unit 105 of the speech recognition system 10 receives the call text from the speech recognition unit 103, it transmits the call text to the terminal 20 that made the speech recognition request (step S515).
[0102] When the UI unit 201 of the terminal 20 receives the call text from the speech recognition system 10, it displays the call details including the call text (step S516). Note that this call detail may be displayed on the search screen or on a screen different from the search screen.
[0103] On the other hand, if it is determined in step S510 above that there is no free resource, the following steps S517 to S519 are executed.
[0104] The speech recognition control unit 102 of the speech recognition system 10 transmits information indicating that speech recognition is impossible to the UI providing unit 105 (step S517).
[0105] When the UI providing unit 105 of the voice recognition system 10 receives information indicating that voice recognition is impossible from the voice recognition control unit 102, it transmits the information to the terminal 20 that made the voice recognition request (step S518).
[0106] When the UI unit 201 of the terminal 20 receives information indicating that voice recognition is impossible from the voice recognition system 10, it displays information indicating that there is no call text (step S519). However, the UI unit 201 may display information other than the call text (for example, call ID, user ID, user name, etc.).
[0107] If voice recognition for the voice data of the call represented by the call data selected by the user is completed, the UI unit 201 of the terminal 20 displays the call details in the same manner as in step S516 above (step S520).
[0108] <Parallel Processing of Voice Recognition> Here, for example, when a real-time call text screen for referring to the call text of the call during the call is displayed, past voice data may also be voice-recognized at the same time. Generally, however, since the voice recognition process takes about the same amount of time as the actual speaking time, it takes a certain amount of time until the user can refer to the call text of the past voice data. Similarly, for example, since it takes a certain amount of time until the call text is created in step S512 of FIG. 9, there may be a waiting time for the user who has viewed the details of the call data.
[0109] Therefore, below, a method for shortening the time until the call text is created by executing voice recognition in parallel will be described. With this method, the voice recognition unit 103 can create the call text in a shorter time.
[0110] For example, when performing speech recognition on certain speech data, in this method, as shown in FIG. 9 for example, first, the speech data is divided into a section called the speech section. Here, the speech section can be detected by a process called voice activity detection (VAD). Then, in this method, as shown in FIG. 9, speech recognition is performed in parallel for each speech section. As a result, since speech recognition is executed in parallel for each speech section, it becomes possible to obtain the call text for the original speech data in a shorter time. Note that since voice activity detection (VAD) can be executed with very few CPU resources compared to the speech recognition process, even if voice activity detection is performed in advance, it hardly affects the resources of the speech recognition system 10.
[0111] <Summary> As described above, in the contact system 1 according to the present embodiment, while preferentially performing speech recognition on the speech data of a call in which a user (operator, supervisor) refers to the call text in real time, for the speech data of a call that is not so, if there is free resource, speech recognition is performed in the background (or during a time period when resources are free such as at night, etc.). Thereby, the resources of the speech recognition system 10 can be used efficiently. For this reason, for example, when some cost occurs according to the multiplicity N of the speech recognition system 10 (for example, when the speech recognition system 10 is realized by a virtual machine on an external cloud server and a cost is incurred according to the number of cores of the CPU of the virtual machine), it becomes possible to reduce that cost.
[0112] The present invention is not limited to the above-described embodiments specifically disclosed, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.
Explanation of Reference Numerals
[0113] 1 Contact Center System 10 Speech Recognition System 20 Terminal 21 Operator Terminal 22 Supervisor Terminal 30 Telephone 40 PBX 50 NW Switch 60 Customer Terminal 70 Communication Network 101 Recording Unit 102 Voice Recognition Control Unit 103 Voice Recognition Unit 104 Search Unit 105 UI Provisioning Unit 106 Voice Data Storage Unit 107 Call Data Storage Unit 108 Call List Storage Unit 109 Display List Storage Unit 201 UI Unit
Claims
1. An audio recognition control unit configured to determine whether to perform real-time audio recognition on audio data obtained from a voice call; An audio recognition unit configured to perform the audio recognition on the audio data determined to perform real-time audio recognition and create text representing the result of the audio recognition; A UI providing unit configured to display a screen on which the text can be referred to in real time on a terminal connected via a communication network; having; The audio recognition control unit, When the screen is displayed on the terminal, it is configured to determine to perform real-time audio recognition on the audio data that is the source of the text that can be referred to on the screen. An audio recognition system.
2. The audio recognition control unit, Every predetermined time, it determines whether there is free space in the resources for the audio recognition. If it is determined that there is free space in the resources, it is configured to determine to perform audio recognition on the audio data for which it has not been determined to perform real-time audio recognition. The audio recognition unit, It is configured to perform the audio recognition on the audio data for which it has not been determined to perform real-time audio recognition and create text representing the result of the audio recognition. The audio recognition system according to claim 1.
3. The audio recognition control unit, It is configured to randomly or according to a predetermined criterion determine one or more pieces of audio data for which to perform the audio recognition from among the audio data for which it has not been determined to perform real-time audio recognition. The audio recognition unit, It is configured to perform the audio recognition on the determined one or more pieces of audio data and create text representing the result of the audio recognition. The audio recognition system according to claim 2.
4. The voice recognition control unit When it is determined to perform real-time voice recognition on the voice data, it further determines whether there is free space in the resource, When it is determined that there is no free space in the resource, it is configured to determine one or more voice data for which voice recognition is to be aborted from among the voice data for which it has not been determined to perform real-time voice recognition, The voice recognition unit configured to abort voice recognition for one or more voice data for which it has been determined to abort voice recognition, the voice recognition system according to claim 2 or 3.
5. The voice recognition control unit When it is determined that there is no free space in the resource, it is configured to determine one or more voice data for which voice recognition is to be aborted, randomly or according to a predetermined criterion, from among the voice data for which it has not been determined to perform real-time voice recognition, the voice recognition system according to claim 4.
6. The UI providing unit configured to display the screen on either or both of the terminal used by the first user making the voice call or the terminal used by the second user monitoring the voice call of the first user, the voice recognition system according to any one of claims 1 to 5.
7. a storage unit configured to store call data related to the voice call, a search unit configured to search for the call data stored in the storage unit based on search conditions specified by the terminal, and having The voice recognition unit configured to perform voice recognition on the voice data corresponding to the call data when the recognized call data is to be displayed on the terminal and the voice recognition for the voice data corresponding to the call data has not been completed, the voice recognition system according to any one of claims 1 to 6.
8. The voice recognition unit The voice recognition system according to any one of claims 1 to 7, configured to divide the voice data into predetermined utterance section units and perform the voice recognition in parallel for each of the divided utterance section units.
9. A voice recognition control procedure for determining whether to perform voice recognition on voice data acquired from a voice call in real time, A voice recognition procedure for performing the voice recognition on the voice data determined to be subject to real-time voice recognition and creating a text representing the result of the voice recognition, A UI provision procedure for causing a screen on which the text can be referred to in real time to be displayed on a terminal connected via a communication network, which is executed by a computer, wherein the voice recognition control procedure determines to perform voice recognition in real time on the voice data that is the source of the text that can be referred to on the screen when the screen is being displayed on the terminal, a voice recognition method.
10. A program for causing a computer to function as the voice recognition system according to any one of claims 1 to 8.
Citation Information
Patent Citations
Call center system and call center management method
JP2021158413A
Management of speech and audio prompts in multimodal interfaces
US6012030A