Speech recognition system, speech recognition method and program

The voice recognition system optimizes resource usage by selectively performing real-time recognition based on the need for viewing results, addressing inefficiencies in existing systems by prioritizing real-time processing only when the recognition output is required.

JP2025122234APending Publication Date: 2025-08-20NTT TECHNOCROSS CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025094002
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-20

AI Technical Summary

Technical Problem

Existing speech recognition systems waste resources, particularly CPU resources, by performing real-time recognition even when the results are not being viewed, such as when no user interface is launched to check the recognition results.

Method used

A voice recognition system that includes a control unit to determine whether to perform real-time voice recognition based on the need for viewing the recognition results, with real-time recognition prioritized for calls where the text is displayed and background processing used when resources are available.

Benefits of technology

This approach optimizes resource usage by prioritizing real-time recognition only when necessary, reducing waste and improving efficiency, especially in contact centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122234000001_ABST
    Figure 2025122234000001_ABST
Patent Text Reader

Abstract

To increase efficiency of used resources for speech recognition.SOLUTION: A speech recognition system according to one embodiment comprises: a speech recognition control unit that determines whether or not to perform real-time speech recognition on speech data acquired from a speech call; a speech recognition unit that performs the speech recognition on speech data determined to undergo real-time speech recognition and creates texts representing the recognition result; and a UI provision unit that displays a screen on a terminal where the texts can be referenced in real time. The speech recognition control unit determines to perform real-time speech recognition on the speech data that becomes a source for the text viewable on the screen when the screen is displayed on the terminal.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice recognition system, a voice recognition method, and a program. [Background technology]

[0002] Speech recognition systems that target contact centers (also called call centers) and record speech during phone calls and convert it into text in real time have been known for some time (for example, Non-Patent Document 1). In such speech recognition systems, speech is generally recorded and speech recognition is performed for all calls made at the contact center. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] ForeSight Voice Mining, Internet<URL:https: / / www.ntt-tx.co.jp / products / foresight_vm / > Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the past, real-time speech recognition was performed even for calls that did not necessarily require real-time speech recognition. For example, real-time speech recognition was performed even when no one was viewing the speech recognition results, such as when the operator had not launched a UI (user interface) to check the speech recognition results. This resulted in wasted resources (especially CPU (Central Processing Unit) resources, etc.).

[0005] An embodiment of the present invention has been made in view of the above points, and aims to improve the efficiency of resources used in speech recognition. [Means for solving the problem]

[0006] In order to achieve the above object, a voice recognition system according to one embodiment includes a voice recognition control unit that determines whether or not to perform real-time voice recognition on voice data acquired from a voice call; a voice recognition unit that performs the voice recognition on the voice data that has been determined to be subject to real-time voice recognition and creates text representing the results of the voice recognition; and a UI provision unit that displays on a terminal a screen on which the text can be viewed in real time, wherein, when the screen is displayed on the terminal, the voice recognition control unit determines to perform real-time voice recognition on the voice data that is the source of the text that can be viewed on the screen. [Effects of the Invention]

[0007] The resources used for voice recognition can be made more efficient. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a contact center system according to an embodiment of the present invention. [Figure 2] FIG. 10 is a diagram illustrating an example of a real-time call text screen. [Figure 3] FIG. 2 is a diagram illustrating an example of the functional configuration of a voice recognition system and a terminal according to the present embodiment. [Figure 4] FIG. 10 is a sequence diagram illustrating an example of a display start process of a real-time call text screen according to the present embodiment. [Figure 5] FIG. 10 is a sequence diagram illustrating an example of a display termination process of the real-time call text screen according to the embodiment. [Figure 6] FIG. 4 is a sequence diagram showing an example of processing from the start of a call to the end of a call according to the present embodiment. [Figure 7] FIG. 10 is a sequence diagram illustrating an example of background speech recognition processing according to the present embodiment. [Figure 8] FIG. 10 is a sequence diagram illustrating an example of a search process according to the present embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of parallel processing of speech recognition. DETAILED DESCRIPTION OF THE INVENTION

[0009] An embodiment of the present invention will be described below. In this embodiment, a contact center system 1 will be described that is targeted at a contact center and can improve the efficiency of resources used for speech recognition (especially CPU resources, etc.) for speech recorded from a call made by an operator. However, the contact center is just one example, and the present invention can be similarly applied to a location other than a contact center, for example, a case where resources used for speech recognition are improved for speech recorded from a call made by a person working in an office, etc. More generally, the present invention can be similarly applied to a case where resources used for speech recognition are improved for speech recorded from a call made by the person working in an office, etc.

[0010] <Overall configuration of contact center system 1> An example of the overall configuration of a contact center system 1 according to this embodiment is shown in Fig. 1. As shown in Fig. 1, the contact center system 1 according to this embodiment includes a voice recognition system 10, multiple terminals 20, multiple telephones 30, a PBX (Private Branch eXchange) 40, a NW switch 50, and a customer terminal 60. Here, the voice recognition system 10, the terminals 20, the telephones 30, the PBX 40, and the NW switch 50 are installed in a contact center environment E, which is the system environment of the contact center. Note that the contact center environment E is not limited to a system environment within the same building, and may be, for example, a system environment within multiple geographically separated buildings.

[0011] The speech recognition system 10 records the voice of a call between an operator and a customer using packets (voice packets) transmitted from the NW switch 50. The speech recognition system 10 also performs speech recognition on the recorded voice and converts it into text (hereinafter also referred to as "call text"). At this time, if the operator or supervisor refers to the call text in real time, the speech recognition system 10 performs speech recognition on the voice of the call between the operator and the customer in real time, but does not perform the speech recognition in real time otherwise. Note that a supervisor is, for example, a person who monitors an operator's calls and supports the operator's telephone answering work when a problem is likely to occur or at the operator's request. Usually, calls of several to a dozen operators are monitored by one supervisor.

[0012] Hereinafter, the screen on which an operator or supervisor can refer to the call text in real time will be referred to as the “real-time call text screen.” The real-time call text screen displays the call text, which is the result of real-time speech recognition, in real time.

[0013] The terminal 20 is a terminal of various types, such as a PC (personal computer) used by an operator or a supervisor. Hereinafter, the terminal 20 used by an operator will be referred to as an "operator terminal 21," and the terminal 20 used by a supervisor will be referred to as a "supervisor terminal 22."

[0014] The telephone 30 is an IP (Internet Protocol) telephone (such as a fixed IP telephone or a mobile IP telephone) used by an operator. Generally, one operator terminal 21 and one telephone 30 are installed at an operator's desk.

[0015] The PBX 40 is a telephone exchange (IP-PBX) and is connected to a communication network 70 including a Voice over Internet Protocol (VoIP) network and a Public Switched Telephone Network (PSTN).

[0016] The NW switch 50 relays packets between the telephone 30 and the PBX 40 , and also captures the packets and transmits them to the voice recognition system 10 .

[0017] The customer terminal 60 is a variety of terminals used by customers, such as a smartphone, a mobile phone, or a landline phone.

[0018] The overall configuration of the contact center system 1 shown in Fig. 1 is an example, and other configurations are also possible. For example, in the example shown in Fig. 1, the PBX 40 is an on-premise telephone exchange, but it may also be a telephone exchange implemented by a cloud service. Also, for example, the voice recognition system 10 may be implemented by a single server and called a voice recognition device. Furthermore, if the operator terminal 21 also functions as an IP telephone, the operator terminal 21 and the telephone 30 may be integrated into one unit.

[0019] <Real-time call text screen> An example of the real-time call text screen is shown in Fig. 2. The real-time call text screen 1000 shown in Fig. 2 includes a real-time call text display field 1100, and each time speech recognition is performed in real time by the speech recognition system 10, the call text obtained by the speech recognition is displayed in real time in the real-time call text display field 1100 (that is, the call text obtained by the speech recognition is immediately displayed in the real-time call text display field 1100).

[0020] For example, in the example shown in FIG. 2, call texts 1101 to 1106 are displayed in a real-time call text display field 1100.

[0021] This allows the operator or supervisor to check the conversation between the operator and the customer currently on the call in real time by referring to the real-time call text screen.

[0022] <Functional configuration of the speech recognition system 10 and the terminal 20> FIG. 3 shows an example of the functional configuration of the speech recognition system 10 and the terminal 20 according to this embodiment.

[0023] <Voice Recognition System 10> As shown in FIG. 3 , the speech recognition system 10 according to this embodiment includes a recording unit 101, a speech recognition control unit 102, a speech recognition unit 103, a search unit 104, and a UI providing unit 105. These units are realized, for example, by a processor such as a CPU executing one or more programs installed in the speech recognition system 10. The speech recognition system 10 according to this embodiment also includes a speech data storage unit 106, a call data storage unit 107, a call list storage unit 108, and a display list storage unit 109. These storage units are realized, for example, by auxiliary storage devices such as HDDs (Hard Disk Drives) and SSDs (Solid State Drives). Note that at least some of these storage units may be realized, for example, by storage devices connected to the speech recognition system 10 via a communication network.

[0024] The recording unit 101 records the voice data contained in the voice packet transmitted from the NW switch 50. That is, the recording unit 101 stores the voice data contained in the voice packet in association with a call ID in the voice data storage unit 106. The call ID is information that uniquely identifies a call between an operator and a customer.

[0025] Furthermore, when a call between an operator and a customer starts, the recording unit 101 adds a pair of the user ID of the operator making the call and the call ID of that call to the call list. Furthermore, when the call ends, the recording unit 101 deletes the pair of the user ID of the operator making the call and the call ID of that call from the call list. Here, the call list is a list that stores pairs of the user ID of the operator currently on the call and the call ID of that call. Note that the user ID is information that uniquely identifies an operator (and a supervisor).

[0026] The voice recognition control unit 102 controls whether or not to perform real-time voice recognition (i.e., instantaneous voice recognition) on calls between an operator and a customer. That is, for calls in which call text is displayed in real time on the real-time call text screen, the voice recognition control unit 102 performs real-time voice recognition on the voice of the call, while for other calls, the voice recognition control unit 102 controls the voice recognition to be performed in the background at some timing rather than in real time. Furthermore, when performing real-time voice recognition on the voice of a new call, if there are insufficient CPU resources, the voice recognition control unit 102 also controls to stop part or all of the background voice recognition and give priority to real-time voice recognition.

[0027] The voice recognition unit 103 performs voice recognition on the voice data to create a call text in accordance with the control of the voice recognition control unit 102. The voice recognition unit 103 also creates call data including at least a call ID and a call text, and stores the call data in the call data storage unit 107.

[0028] The search unit 104 searches for call data stored in the call data storage unit 107 based on the search conditions received from the UI providing unit 105 .

[0029] The UI providing unit 105 provides information (hereinafter also referred to as UI information) for displaying UIs (user interfaces) of various screens (for example, a real-time call text screen, a search screen for the user to specify the above search conditions, etc.) on the terminal 20. The UI information may be any information necessary for displaying a screen, and examples thereof include screen definition information in which a screen is defined by HTML (Hypertext Markup Language) or the like.

[0030] Furthermore, when the UI providing unit 105 receives a display request for a real-time call text screen from the terminal 20, it adds the pair of user IDs included in the display request to the display list. Furthermore, when the display of the real-time call text screen ends, the UI providing unit 105 deletes the pair of user IDs included in the end notification from the display list. Here, the display list is a list that stores pairs of the user ID of the operator making the call whose call text is displayed in real time on the real-time call text screen and the user ID of the user (operator or supervisor) of the terminal 20 on which the real-time call text screen is displayed.

[0031] The audio data storage unit 106 stores the audio data recorded by the recording unit 101.

[0032] The call data storage unit 107 stores call data. The call data includes at least a call ID and a call text, but may also include various other information such as a caller's telephone number and a destination telephone number related to the call of the call ID, the user ID of the operator who made the call, the call start time and the call end time of the call, etc.

[0033] The call list storage unit 108 stores a call list in which pairs of the user ID of the operator currently in a call and the call ID of that call are stored.

[0034] The display list storage unit 109 stores a display list that stores pairs of the user ID of the operator making the call whose call text is displayed in real time on the real-time call text screen and the user ID of the user of the terminal 20 on which the real-time call text screen is displayed.

[0035] <Terminal 20> 3, the terminal 20 according to this embodiment includes a UI unit 201. The UI unit 201 is realized, for example, by processing in which one or more programs installed in the terminal 20 are executed by a processor such as a CPU.

[0036] The UI unit 201 displays various screens (for example, a real-time call text screen, a search screen, etc.) on a display or the like based on UI information provided from the UI providing unit 105 of the speech recognition system 10. The UI unit 201 also accepts various operations on the screens displayed on the display or the like.

[0037] <Processing of Contact Center System 1> Various processes executed by the contact center system 1 according to this embodiment will be described below.

[0038] <<Real-time call text screen display start process>> The display start process of the real-time call text screen according to this embodiment will be described with reference to Fig. 4. The following describes a case where a certain user (operator or supervisor) displays the real-time call text screen on the display of his / her own terminal 20.

[0039] If the real-time call text screen is not being displayed, it can be displayed at any time (i.e., this process can be started at any time). Therefore, for example, if a real-time call text screen on which the call text of a certain operator's call is displayed in real time is to be displayed on the terminal 20, the user (the operator himself or a supervisor monitoring the operator's calls) can display the real-time call text screen before the start of the call, or can display the real-time call text screen during the call.

[0040] First, the UI unit 201 of the terminal 20 transmits a display request for a real-time call text screen to the speech recognition system 10 in response to an operation for displaying the real-time call text screen (step S101). Here, the display request includes the user ID of the operator who is to have the call text displayed in real time on the real-time call text screen (hereinafter also referred to as the display target user ID), and the user ID of the user who uses the terminal 20 that transmitted the display request (hereinafter also referred to as the display user ID). Note that if the terminal 20 is an operator terminal 21, the display target user ID and the display user ID are the user ID of the operator who uses the operator terminal 21. On the other hand, if the terminal 20 is a supervisor terminal 22, the display target user ID is the user ID of a certain operator who is monitored by the supervisor terminal 22, and the display user ID is the user ID of the supervisor who uses the supervisor terminal 22.

[0041] When the UI providing unit 105 of the voice recognition system 10 receives a display request for a real-time call text screen, the UI providing unit 105 adds the display target user ID and the display user ID included in the display request to the display list (step S102).

[0042] Next, the UI providing unit 105 of the voice recognition system 10 transmits UI information of the real-time call text screen to the terminal 20 (step S103).

[0043] When the UI unit 201 of the terminal 20 receives the UI information of the real-time call text screen, the UI unit 201 displays the real-time call text screen on the display based on the UI (step S104).

[0044] <<Real-time call text screen display termination process>> The display termination process of the real-time call text screen according to this embodiment will be described with reference to Fig. 5. The following describes a case where a certain user (operator or supervisor) terminates the display of the real-time call text screen displayed on the display of his / her own terminal 20.

[0045] When the real-time call text screen is displayed, the display of the real-time call text screen can be ended at any timing (that is, this process can be started at any timing). Therefore, for example, when a real-time call text screen on which the call text of a call of a certain operator is displayed in real time is displayed on the terminal 20, the user (the operator himself or a supervisor monitoring the operator's calls) can end the display of the real-time call text screen during the call, or can end the display of the real-time call text screen after the call has ended.

[0046] First, the UI unit 201 of the terminal 20 ends the display of the real-time call text screen in response to an operation for ending the display of the real-time call text screen (step S201).

[0047] Next, the UI unit 201 of the terminal 20 transmits a display end notification to the speech recognition system 10 (step S202). Here, the display end notification includes the display target user ID and the display user ID. If the terminal 20 is an operator terminal 21, the display target user ID and the display user ID are the user ID of an operator who uses the operator terminal 21. On the other hand, if the terminal 20 is a supervisor terminal 22, the display target user ID is the user ID of a certain operator who was supervising at the supervisor terminal 22, and the display user ID is the user ID of a supervisor who uses the supervisor terminal 22.

[0048] When receiving the display end notification, the UI providing unit 105 of the voice recognition system 10 deletes the display target user ID and the display user ID included in the display end notification from the display list (step S203).

[0049] <Processing from the start of a call to the end of a call> The process from the start of a call to the end of a call according to this embodiment will be described with reference to Fig. 6. The process from the start of a call to the end of a call by a certain operator will be described below.

[0050] First, the recording unit 101 of the voice recognition system 10 receives a call start packet from the NW switch 50 (step S301).

[0051] Next, the recording unit 101 of the speech recognition system 10 adds the user ID included in the call start packet (hereinafter also referred to as the active user ID) and the call ID of the call that has started to the call list (step S302). Note that the call ID is generated arbitrarily by the recording unit 101, but since one operator can only handle one call at a time, for example, the call ID may be generated by adding the call start date and time, etc. to the active user ID.

[0052] The following steps S303 to S315 are repeatedly executed during the call (that is, until the recording unit 101 receives a call end packet). Steps S303 to S315 in one repetition will be described below.

[0053] The recording unit 101 of the voice recognition system 10 receives a voice packet from the NW switch 50 (step S303). Here, the voice packet includes voice data and a user ID (a user ID during a call), and the recording unit 101 identifies a call ID corresponding to the user ID during a call from the call list, and then stores the voice data in the voice data storage unit 106 in association with the identified call ID.

[0054] The recording unit 101 of the voice recognition system 10 transmits the active user ID included in the voice packet received from the NW switch 50 to the voice recognition control unit 102 (step S304).

[0055] When the voice recognition control unit 102 of the voice recognition system 10 receives the active user ID, it determines whether or not it is necessary to perform real-time voice recognition on the voice data of the call ID corresponding to the active user ID in the call list (step S305). Specifically, the voice recognition control unit 102 determines whether or not the active user ID is included as a display target user ID in the display list. If the active user ID is included as a display target user ID in the display list, the voice recognition control unit 102 determines that it is necessary to perform real-time voice recognition on the voice data of the call ID corresponding to the active user ID in the call list; otherwise, it determines that it is not necessary to perform real-time voice recognition on the voice data. Note that if the active user ID is included as a display target user ID in the display list, this means that the call text of the call being made by the operator of the active user ID is referenced in real time on the real-time call text screen.

[0056] If it is determined in step S305 above that real-time speech recognition is necessary, the following steps S306 to S315 are executed.

[0057] The speech recognition control unit 102 of the speech recognition system 10 determines whether there is free space in the resources available for speech recognition (especially CPU resources, etc.) (step S306). Here, the resources available for speech recognition are often represented by an index value called multiplicity, which represents the number of speech data that can be recognized simultaneously. For example, if the multiplicity is N, it means that N pieces of speech data can be recognized simultaneously. Therefore, the speech recognition control unit 102 can, for example, determine that there is free space in the resources if the number of speech data currently being recognized simultaneously is n and the multiplicity is N and n < N, and determine that there is no free space in the resources otherwise.

[0058] If it is determined in step S306 above that there is no free space in the resources, the following steps S307 to S309 are executed.

[0059] The speech recognition control unit 102 of the speech recognition system 10 determines the speech data for which speech recognition is to be aborted from the speech data stored in the speech data storage unit 106 according to the following procedures 1 to 3 (step S307).

[0060] Procedure 1: The speech recognition control unit 102 identifies the speech data currently being recognized among the speech data stored in the speech data storage unit 106.

[0061] Procedure 2: Next, the speech recognition control unit 102 identifies the speech data other than the speech data during real-time speech recognition among the speech data identified in Procedure 1. Here, the speech data during real-time speech recognition can be identified by identifying the in-call user ID included as the display target user ID in the display list and then identifying the call ID corresponding to these in-call user IDs from the call list, and then identifying the speech data associated with these call IDs.

[0062] Step 3: Then, the voice recognition control unit 102 determines one or more voice data from the voice data identified in step 2 as voice data for which voice recognition is to be stopped. The voice data for which voice recognition is to be stopped may be one or more. The voice data may be determined randomly from the voice data identified in step 2, or may be determined according to some criteria. Examples of such criteria include giving priority to the voice data that has been in use for a shorter (or longer) time since the start of voice recognition, giving priority to the voice data of a call made by a specific operator (or an operator belonging to a specific group), determining according to a round robin method, etc.

[0063] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID associated with the voice data for which it has been determined in step S307 that voice recognition is to be stopped to the voice recognition unit 103 (step S308).

[0064] The voice recognition unit 103 of the voice recognition system 10 stops voice recognition of the voice data associated with the call ID received from the voice recognition control unit 102 (step S309). This makes available resources available for voice recognition.

[0065] If it is determined in step S306 above that there are available resources or following step S309 above, the speech recognition control unit 102 of the speech recognition system 10 identifies from the call list a call ID corresponding to the active user ID sent from the recording unit 101 in step S304 above, and sends the identified call ID and the active user ID to the speech recognition unit 103 (step S310).

[0066] The voice recognition unit 103 of the voice recognition system 10 performs voice recognition on the voice data associated with the call ID received from the voice recognition control unit 102 (step S311). As a result, a call text is created as a result of the voice recognition performed on the voice data.

[0067] For example, during a call, a real-time call text screen may be displayed on a certain terminal 20 to refer to the call text of that call. In this case, there may be no call text until the real-time call text screen is displayed. For example, at time t s For calls that start at a certain time t (>t s ) displays the real-time call text screen to view the call text for that call, and at time t s In this case, in step S311, the speech recognition unit 103 recognizes not only the speech data after time t but also the past speech data (i.e., for example, the speech data after time t s The speech data from t to t may also be simultaneously recognized.

[0068] The voice recognition unit 103 of the voice recognition system 10 transmits the call text created in the above step S311 and the active user ID received from the voice recognition control unit 102 in the above step S310 to the UI provision unit 105 (step S312).

[0069] Furthermore, the speech recognition unit 103 of the speech recognition system 10 associates the call text created in step S311 with the call ID and stores it as call data in the call data storage unit 107 (step S313). At this time, various information such as the ID of the user currently on the call may be included in the call data.

[0070] When the UI providing unit 312 of the speech recognition system 10 receives the call text and the user ID of the user currently on the call, it identifies from the display list the display user ID corresponding to the display target user ID that matches the currently on the call user ID, and sends the call text to the terminal 20 of the identified display user ID (step S314).

[0071] When the UI unit 201 of the terminal 20 receives the call text from the voice recognition system 10, it displays the call text on the real-time call text screen (step S315). As a result, the call text is displayed in real time on the real-time call text screen.

[0072] When the call termination packet is transmitted from the NW switch 50, the recording unit 101 of the voice recognition system 10 receives the call termination packet from the NW switch 50 (step S316).

[0073] Then, the recording unit 101 of the voice recognition system 10 deletes the active user ID that matches the user ID included in the call end packet and the corresponding call ID from the call list (step S317).

[0074] <Background speech recognition processing> The background speech recognition process according to this embodiment will be described with reference to FIG. 7. This background speech recognition process is a process for performing speech recognition on speech data other than the speech data that was the target of real-time speech recognition, and is repeatedly executed at predetermined intervals (e.g., every 10 minutes) in the background of the above-mentioned "process for starting display of real-time call text screen," "process for ending display of real-time call text screen," and "processing from the start of a call to the end of a call." However, the time interval for repeating the background speech recognition process may vary depending on, for example, the time period. For example, the repetition time interval may be longer during the daytime when call volume is high to perform more real-time speech recognition, and shorter during the nighttime when call volume is low to perform more background speech recognition. Alternatively, the background speech recognition process may not be executed during the daytime when call volume is high to perform more real-time speech recognition.

[0075] First, the voice recognition control unit 102 of the voice recognition system 10 determines whether or not there is free space in resources (especially CPU resources, etc.) available for voice recognition, similarly to step S306 in FIG. 6 (step S401).

[0076] If it is determined in step S401 above that there are no available resources, the following steps S402 to S404 are executed.

[0077] The voice recognition control unit 102 of the voice recognition system 10 determines voice data to be recognized from the voice data stored in the voice data storage unit 106 according to the following steps 11 and 12 (step S402).

[0078] Step 11: The voice recognition control unit 102 identifies voice data that is not currently being recognized from the voice data stored in the voice data storage unit 106.

[0079] Step 12: Then, the voice recognition control unit 102 determines one or more voice data from the voice data identified in step 11 as the voice data to be recognized. Note that the voice data to be recognized may be one or more depending on the availability of resources available for voice recognition. The voice data may be determined randomly from the voice data identified in step 11, or may be determined according to some criteria. Examples of such criteria include giving priority to the voice data that has been in use for a long (or short) time since the start of voice recognition, giving priority to the voice data of a call made by a specific operator (or an operator belonging to a specific group), determining according to a round robin method, etc.

[0080] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID associated with the voice data determined to be voice-recognized in step S402 above to the voice recognition unit 103 (step S403).

[0081] The voice recognition unit 103 of the voice recognition system 10 performs voice recognition on the voice data associated with the call ID received from the voice recognition control unit 102 (step S404). As a result, a call text is created as a result of the voice recognition performed on the voice data.

[0082] The speech recognition unit 103 of the speech recognition system 10 associates the call text created in step S404 with the call ID and stores it as call data in the call data storage unit 107 (step S405). At this time, various information such as the user ID of the operator who made the call with this call ID may be included in the call data.

[0083] <<Search processing>> The search process according to this embodiment will be described with reference to Fig. 8. In the following, a case where a certain user (operator or supervisor) searches for call data using his / her own terminal 20 will be described.

[0084] The search for call data can be performed at any time (that is, the execution of this process can be started at any time).

[0085] The UI unit 201 of the terminal 20 transmits a search request including search conditions specified by the user to the speech recognition system 10 (step S501). Any conditions for searching call data can be specified as the search conditions, such as a user ID, a call start date and time, a call end date and time, and a call duration. The user can specify the search conditions on a search screen for specifying search conditions, for example.

[0086] When the UI providing unit 105 of the voice recognition system 10 receives a search request from the terminal 20, it transmits the search request to the searching unit 104 (step S502).

[0087] When the search unit 104 of the voice recognition system 10 receives the search request from the UI providing unit 105, it searches the call data stored in the call data storage unit 107 based on the search conditions included in the search request (step S503).

[0088] The search unit 104 of the voice recognition system 10 transmits the search result in the above step S503 to the UI providing unit 105 (step S504). Note that the search result includes, for example, the call data searched in the above step S503.

[0089] When the UI providing unit 105 of the voice recognition system 10 receives the search results from the searching unit 104, the UI providing unit 105 transmits the search results to the terminal 20 (step S505).

[0090] When the UI unit 201 of the terminal 20 receives the search results from the speech recognition system 10, it displays a search result list, which is a list of call data included in the search results (step S506). The user can select call data that the user desires to display in detail from this search list. Note that this search result list may be displayed on the search screen or on a screen different from the search screen.

[0091] The UI unit 201 of the terminal 20 accepts the selection of call data to be displayed in detail from the list of search results (step S507).

[0092] Here, if speech recognition has been completed for the voice data of the call represented by the call data selected by the user, the call data includes the call text for the entire call. On the other hand, if speech recognition has not been completed for the voice data of the call represented by the call data selected by the user, the call data does not include the call text, or only includes the call text for part of the call. Therefore, if speech recognition has not been completed for the voice data of the call represented by the call data selected by the user, the following steps S508 to S519 are executed, and if not, the following step S520 is executed. Whether the call text is only part of the call can be determined, for example, from the call duration, etc.

[0093] The UI unit 201 of the terminal 20 transmits a voice recognition request to the voice recognition system 10 (step S508). Here, the voice recognition request includes the call ID of the call data selected by the user.

[0094] When the UI providing unit 105 of the voice recognition system 10 receives the voice recognition request from the terminal 20, it transmits the voice recognition request to the voice recognition control unit 102 (step S509).

[0095] The voice recognition control unit 102 of the voice recognition system 10 determines whether or not there is free space in resources (particularly CPU resources, etc.) available for voice recognition, similarly to step S306 in FIG. 6 (step S510).

[0096] If it is determined in step S510 above that there is available resource, the following steps S511 to S516 are executed.

[0097] The voice recognition control unit 102 of the voice recognition system 10 transmits the call ID included in the voice recognition request received from the UI providing unit 105 to the voice recognition unit 103 (step S511).

[0098] The voice recognition unit 103 of the voice recognition system 10 performs voice recognition on the voice data associated with the call ID received from the voice recognition control unit 102 (step S512). As a result, a call text is created as a result of the voice recognition performed on the voice data.

[0099] The voice recognition unit 103 of the voice recognition system 10 transmits the call text created in step S512 to the UI provision unit 105 (step S513).

[0100] Furthermore, the speech recognition unit 103 of the speech recognition system 10 associates the call text created in step S512 with the call ID and stores it as call data in the call data storage unit 107 (step S514). At this time, various information such as the ID of the user currently on the call may be included in the call data.

[0101] When the UI providing unit 105 of the voice recognition system 10 receives the call text from the voice recognition unit 103, the UI providing unit 105 transmits the call text to the terminal 20 that has made the voice recognition request (step S515).

[0102] When the UI unit 201 of the terminal 20 receives the call text from the voice recognition system 10, it displays the call details including the call text (step S516). Note that the call details may be displayed on the search screen or on a screen different from the search screen.

[0103] On the other hand, if it is determined in step S510 above that there are no available resources, the following steps S517 to S519 are executed.

[0104] The voice recognition control unit 102 of the voice recognition system 10 transmits information indicating that voice recognition is not possible to the UI providing unit 105 (step S517).

[0105] When the UI providing unit 105 of the voice recognition system 10 receives the information indicating that voice recognition is not possible from the voice recognition control unit 102, the UI providing unit 105 transmits the information to the terminal 20 that has made the voice recognition request (step S518).

[0106] When receiving the information indicating that speech recognition is not possible from the speech recognition system 10, the UI unit 201 of the terminal 20 displays information indicating that there is no call text (step S519). However, the UI unit 201 may display information other than the call text (for example, a call ID, a user ID, a user name, etc.).

[0107] If the voice recognition for the voice data of the call represented by the call data selected by the user has been completed, the UI unit 201 of the terminal 20 displays the call details in the same manner as in step S516 above (step S520).

[0108] <Parallel processing of speech recognition> Here, for example, when a real-time call text screen for referencing the call text of a call is displayed during a call, past voice data may also be simultaneously recognized. However, since the voice recognition process generally takes approximately the same amount of time as the actual speaking time, it takes a certain amount of time before the user can refer to the call text of the past voice data. Similarly, since it takes a certain amount of time for the call text to be created in step S512 of Fig. 9, the user who has displayed the detailed call data may experience a waiting time.

[0109] Therefore, the following describes a method for reducing the time required to create a call text by performing speech recognition in parallel. This method enables the speech recognition unit 103 to create a call text in a shorter time.

[0110] For example, when performing speech recognition on certain speech data, in this method, the speech data is first divided into sections called speech sections, as shown in FIG. 9. Here, speech sections can be detected by a process called voice activity detection (VAD). Then, in this method, speech recognition is performed in parallel for each speech section, as shown in FIG. 9. This allows speech recognition to be performed in parallel for each speech section, making it possible to obtain a call text for the original speech data in a shorter time. Note that, since voice activity detection (VAD) can be performed with significantly less CPU resource than speech recognition processing, performing speech activity detection in advance has almost no impact on the resources of the speech recognition system 10.

[0111] <Summary> As described above, in the contact system 1 according to this embodiment, speech recognition is performed preferentially on speech data of calls in which the user (operator, supervisor) refers to the call text in real time, while speech recognition is performed in the background (or during times when resources are available, such as at night) for speech data of other calls if resources are available. This allows the resources of the speech recognition system 10 to be used efficiently. Therefore, for example, in cases where some cost is incurred according to the multiplicity N of the speech recognition system 10 (for example, in cases where the speech recognition system 10 is implemented as a virtual machine on an external cloud server and costs are incurred according to the number of CPU cores of the virtual machine), it is possible to reduce the cost.

[0112] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims. [Explanation of symbols]

[0113] 1. Contact Center System 10 Voice Recognition System 20 terminals 21 Operator terminal 22 Supervisor Terminal 30 telephone 40 PBX 50 Network Switches 60 Customer terminals 70 Communication Network 101 Recording Section 102 Voice recognition control unit 103 Voice Recognition Unit 104 Search Section 105 UI provision department 106 Audio data storage unit 107 Call data storage unit 108 Call list storage unit 109 Display List Storage 201 UI section

Claims

1. a voice recognition control unit that determines whether or not to perform voice recognition in real time on voice data acquired from a voice call; a speech recognition unit that performs real-time speech recognition on speech data that has been determined to be speech-recognized in real time, and creates text representing the results of the speech recognition; a UI providing unit that displays a screen on a terminal on which the text can be referenced in real time; and The voice recognition control unit When the screen is displayed on the terminal, the speech recognition system determines to perform real-time speech recognition on speech data that is the source of creating text that can be referenced on the screen.

2. a voice recognition control procedure for determining whether or not to perform voice recognition in real time on voice data acquired from a voice call; a speech recognition procedure for performing real-time speech recognition on speech data determined to be speech-recognized in real time, and creating text representing the results of the speech recognition; a UI providing step of displaying on a terminal a screen on which the text can be referenced in real time; The computer executes The speech recognition control procedure includes: The speech recognition method determines that, when the screen is displayed on the terminal, real-time speech recognition is to be performed on speech data that is the source of creation of text that can be referenced on the screen.

3. a voice recognition control procedure for determining whether or not to perform voice recognition in real time on voice data acquired from a voice call; a speech recognition procedure for performing real-time speech recognition on speech data determined to be speech-recognized in real time, and creating text representing the results of the speech recognition; a UI providing step of displaying on a terminal a screen on which the text can be referenced in real time; on the computer, The speech recognition control procedure includes: When the screen is displayed on the terminal, the program determines to perform real-time speech recognition on the speech data that is the source of the text that can be referenced on the screen.

Citation Information

Patent Citations

  • Call center system and call center management method

    JP2021158413A

  • Translation method and electronic device

    JP2022508789A

  • Management of speech and audio prompts in multimodal interfaces

    US6012030A

  • Voice recognition system

    WO2011074260A1